Tag: Biomedical AI

  • Turning Evidence Into Action: The AEQUITAS Database

    Turning Evidence Into Action: The AEQUITAS Database

    AI tools used in healthcare can carry hidden biases, sometimes performing less accurately for women, ethnic minorities, or other underrepresented groups. But knowing this risk exists in general is one thing. Being able to check whether a specific tool, in a specific clinical area, has documented bias issues is another, and that’s exactly the gap the AEQUITAS Database was built to fill.

    The AEQUITAS Database is a structured, digital repository that systematically collects, organises, and provides access to evidence on gender and racial bias in biomedical AI. Rather than leaving healthcare professionals, researchers, and policymakers to piece together scattered studies and reports, it brings this dispersed evidence together into one coherent, accessible knowledge base.

    The database is publicly accessible through the AEQUITAS project website, and it was built with a user-oriented approach, meaning both technical and non-technical users can navigate it effectively.

    Built for multiple audiences

    The Database was designed to serve several groups at once, each with different needs. For healthcare professionals, it’s a resource for identifying potential biases in AI-based diagnostic and decision-support tools, improving patient-centred care, and reducing the risk of misdiagnosis. For civil society organisations, it supports awareness-raising and advocacy for fair, inclusive healthcare systems. And for researchers and policymakers, it feeds into further scientific investigation, the development of clinical guidelines, and regulatory oversight at European level.

    A practical tool

    The Database is designed to be used, not just consulted occasionally. Before adopting a new AI tool, healthcare professionals can search for entries related to its disease area to check whether bias has been documented in similar tools. When an AI system produces unexpected outputs, the Database can help determine whether it’s a known pattern rather than an isolated anomaly. And when caring for patients from underrepresented groups, it can help clinicians understand whether the tools involved in their care carry documented disparities relevant to that patient’s profile.

    The Database is also a living resource that will grow as new evidence becomes available. The more healthcare professionals, researchers, and organisations engage with it, the more valuable it becomes for everyone.

  • From Principles to Practice: How AEQUITAS Turns Policy into Action

    From Principles to Practice: How AEQUITAS Turns Policy into Action

    AI systems used in healthcare are governed by strict ethical principles and legal requirements, from the EU AI Act to the EU Charter of Fundamental Rights. But knowing the rules isn’t the same as knowing how to apply them. Hospitals and civil society organisations (CSOs) need concrete, operational processes they can actually put into practice, day to day, to make sure these tools are used safely and fairly.

    This is exactly what the AEQUITAS AI Regulatory Model was designed to do. It gives hospitals and CSOs a practical governance framework covering the full lifecycle of AI systems, from procurement and validation to deployment, monitoring, and accountability.

    The model draws inspiration from approaches that have proven effective in other high-risk sectors. It echoes the logic of pharmacovigilance, the systematic tracking of adverse drug reactions, in what AEQUITAS calls an “AI-vigilance” system: a way to report and analyse incidents involving AI-supported clinical decisions, with a particular focus on equity outcomes. It also mirrors the patient safety culture familiar from aviation-inspired approaches in hospitals, built around identifying hazards, reporting incidents, and continuously improving.

    At its core, the model is grounded in the EU Charter of Fundamental Rights (2009), particularly the principles of human dignity, equality and non-discrimination, and the right to healthcare, translating these rights into operational safeguards that follow an AI system throughout its use.

    The five checkpoints

    Rather than treating AI governance as a single box to tick, the model structures oversight around five sequential stages, each one a checkpoint designed to catch potential equity risks before they reach patients. Throughout this process, hospitals act as the main deployers, leading decisions and operational control, while CSOs play an independent oversight role, reviewing documentation, flagging equity-related risks, and conducting audits, without controlling day-to-day clinical decisions themselves.

    These stages follow the natural lifecycle of any AI system in a hospital. It starts with scope and classification, determining an AI tool’s regulatory status before it’s even procured or developed. This sets the direction for everything that follows. Next comes procurement, often the moment hospitals have the greatest leverage, as it’s where suppliers can be required to commit to addressing bias risks throughout a system’s lifecycle. At this stage, CSOs are able to review procurement documentation and flag equity concerns. The third stage, validation, means independently evaluating how a system performs, not just overall, but across different patient groups, before it’s ever used with real patients.

    Once a system passes validation, the deployment stage ensures the right safeguards are in place: trained oversight staff, mechanisms for clinicians to override AI recommendations, and clear communication with patients about when these tools are being used in their care. And because AI systems aren’t static, the final stage, ongoing monitoring, auditing, and accountability, keeps watching for emerging disparities long after a system goes live, with CSOs conducting independent audits and helping escalate persistent issues.

    Building a shared network across Europe

    Beyond the model itself, AEQUITAS is currently establishing a European network of CSOs and public hospitals. This shared governance layer means that insights, incidents, and mitigation strategies identified in one country contribute to a growing, Europe-wide knowledge base, one that benefits patients well beyond the institution where the issue was first spotted.

    Together, these tools turn abstract principles into something hospitals and CSOs can actually act on, closing the gap between good intentions and equitable, everyday care.

    References:

    European Union. (2009). Charter of Fundamental Rights of the European Union. Official Journal of the European Union, C 303/1. https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:12012P/TXT

  • Beyond Ethics: How Europe Is Regulating AI in Healthcare

    Beyond Ethics: How Europe Is Regulating AI in Healthcare

    Knowing that AI can be biased is one thing. Knowing what to do about it, legally and ethically, is another. 

    Medicine’s ethical foundations, extended to AI

    Medicine has long been associated with high ethical standards, from the Hippocratic Oath to the Declarations of Geneva and Helsinki. Unfortunately, the history of medicine includes deeply troubling chapters, from eugenics, forced sterilisation of indigenous and marginalised populations, to non-consensual experimentation on people with disabilities. These historical failures are directly relevant today because the structural injustices behind them shaped decades of medical data collection, and their legacy is exactly what today’s AI systems risk inheriting when trained on that same data.

    Biomedical ethics is built around four core ethical principles: autonomy (respecting a person’s right to make their own decisions), non-maleficence (avoiding harm), beneficence (actively promoting wellbeing), and justice (distributing benefits and risks fairly, especially given known disparities based on race, gender, and social status) (Beauchamp & Childress, 2019).

    As AI entered clinical practice, ethicists added a fifth principle specifically for these technologies: explicability, the idea that the reasoning behind an AI system’s decisions must be understandable and open to scrutiny, not simply accepted at face value (Floridi et al., 2018). 

    The EU’s legal response

    The European Union has built a multi-layered legal framework specifically to address these risks. At its centre is the AI Act, which classifies many healthcare applications, including diagnostic tools and clinical decision-support systems, as “high-risk.” This means they’re subject to strict requirements: providers must actively identify and mitigate risks of discriminatory outcomes, ensure meaningful human oversight, and maintain accuracy and transparency throughout the system’s lifecycle (EU AI Act, 2024).

    Hospitals that deploy these systems have specific legal obligations too, including using AI in line with provider instructions, assigning competent human oversight, maintaining logs of system operation, and reporting serious incidents, particularly where there’s a risk of unequal outcomes across patient groups.

    The AI Act doesn’t stand alone. It works alongside the Medical Device Regulation (MDR; European Union, 2017a) and In Vitro Diagnostic Regulation (IVDR; European Union, 2017b), which govern the safety of medical technologies, the GDPR (European Union, 2016), which protects the sensitive health data these systems rely on, and the emerging European Health Data Space (EHDS; European Union, 2025), designed to enable more representative datasets for future AI development. Together, they form a legal ecosystem, all ultimately grounded in the EU Charter of Fundamental Rights (2009), particularly its guarantees of non-discrimination, equality, and the right to healthcare.

    All these policies translate into real responsibilities for hospitals and healthcare professionals, from understanding a tool’s limitations and knowing when to override its recommendations, to being transparent with patients about when and how these systems are used in their care.

    Beyond raising awareness, the AEQUITAS project has developed a practical AI Regulatory Model that translates these legal and ethical principles into concrete steps hospitals and healthcare professionals can actually apply, turning complex EU legislation into something usable in everyday clinical practice.

    References:

    Beauchamp, T. L., & Childress, J. F. (2019). Principles of biomedical ethics (8th ed).

    European Commission, Timeline for Implementation of the EU AI Act, Brussels: European Commission, 2024.

    European Union. (2009). Charter of Fundamental Rights of the European Union. Official Journal of the European Union, C 303/1. https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:12012P/TXT

    European Union, Regulation (EU) 2016/679 on General Data Protection Regulation (GDPR), Brussels: EU, 2016.

    European Union, Regulation (EU) 2017/745 on Medical Devices (MDR), Official Journal of the European Union, L117, Brussels: European Union, 2017.

    European Union, Regulation (EU) 2017/746 on In Vitro Diagnostic Medical Devices (IVDR), Official Journal of the European Union, L117, Brussels: European Union, 2017.

    European Union, Regulation (EU) 2025/327 on the European Health Data Space (EHDS), Brussels: European Union, 2025.

    Floridi, L., Cowls, J., Beltrametti, M., Chatila, R., Chazerand, P., Dignum, V., Luetge, C., Madelin, R., Pagallo, U., Rossi, F., Schafer, B., Valcke, P., & Vayena, E. (2018). AI4People—An Ethical Framework for a Good AI Society: Opportunities, Risks, Principles, and Recommendations. Minds and Machines, 28(4), 689–707. https://doi.org/10.1007/s11023-018-9482-5

  • When “Neutral” Data Isn’t: Gender and Racial Bias in Biomedical AI

    When “Neutral” Data Isn’t: Gender and Racial Bias in Biomedical AI

    Bias in AI isn’t abstract. It shows up in real diagnoses, real treatment decisions, and real patients, particularly those who already face multiple layers of disadvantage. Discrimination in AI is an intersectional issue, overlapping with age, gender, race, sexual orientation, and health status, among other dimensions of identity. The patients most at risk of being harmed by biased AI tools are often those who already face the most barriers to quality care.

    Decades of biased medical data

    The biases we see in AI didn’t come from nowhere. They’re a direct consequence of medical data collected over decades, data shaped by systemic discrimination and structural inequalities in healthcare research and practice (Cirillo et al., 2020; Cross et al., 2024).

    For much of medical history, clinical and experimental studies focused predominantly on male participants. This is what is known as the gender health gap. In 2020, only 5% of global health research funding went to women’s health research. This was split into 4% for women’s cancers and 1% for all other women-specific health conditions, with 25% of that further limited to fertility research (Nature Reviews Bioengineering, 2024). 

    It’s worth noting this isn’t a one-way problem: in some areas, such as depression, men are underrepresented in clinical data, largely because they’re less likely to seek care, report symptoms, or receive a diagnosis (Smith et al., 2018).

    Racial bias is equally well documented

    Racial bias in medicine is also extensively documented, particularly in the United States where research shows that, for instance, Black patients and other minority groups receive fewer medical procedures, lower rates of surgical intervention, and fewer referrals to specialists than white patients, regardless of clinical need (Bowser, 2001; Williams & Wyatt, 2015). When AI systems are trained on data that reflects decades of existing inequalities in healthcare access and treatment, they risk encoding and perpetuating those inequalities at a much larger scale. While much of the evidence comes from the US, the structural conditions producing racial bias, including socioeconomic inequalities, barriers to access and underrepresentation in research, are present across Europe too.

    LGBTQIA+ patients face their own layer of risk

    It is also important to recognise the specific situation of LGBTQIA+ individuals, who experience discrimination in healthcare and are subject to stereotypes that affect the care they receive. These social and cultural factors perpetuate discrimination and have a measurable impact on health and healthcare. Research has shown that 16% of LGBTQIA+ individuals report discrimination in healthcare encounters, and 18% avoid seeking care altogether due to fear of mistreatment (Chang et al., 2025). Another large-scale study found that in emergency department scenarios, AI recommendations for LGBTQIA+ patients included mental health interventions six to seven times more often than was clinically appropriate (Chang et al., 2025).

    How this plays out in three key areas

    In cardiovascular care, women’s symptoms often differ from the “classic” presentation described in medical textbooks, itself largely based on male patient data (Fatunde et al., 2025). As a result, women are offered fewer diagnostic tests, less medication, and fewer specialist referrals (Al Hamid et al., 2024). Racial bias compounds the problem: pulse oximeters, devices routinely used to measure blood oxygen saturation, have been shown to produce less accurate readings for patients with darker skin tones (Sjoding et al., 2020), a bias that then feeds directly into AI-based triage systems.

    In diabetes care, racial bias is extensively documented. Studies have shown that African American patients present systematically higher A1c values than white patients with the same average blood glucose (Karter et al., 2023). If an AI system uses A1c as a proxy for glycaemic control without accounting for this, it risks producing incorrect diagnoses for Black patients.

    In depression, both gender and racial bias carry significant weight, especially in tools using natural language processing for screening. Men and women tend to express psychological distress differently (Pennebaker et al., 2003), so a system trained predominantly on one gender’s language patterns may screen inaccurately for the other. Most mental health AI tools also still operate on binary gender assumptions, excluding non-binary, transgender, and gender non-conforming individuals from both the data and the populations these tools are meant to serve (Hafner et al., 2024).

    Why this matters

    These biases translate into delayed diagnoses, inappropriate treatments, and unequal care, every day, for real patients. Understanding these patterns is exactly what equips healthcare professionals to ask better questions, challenge assumptions, and advocate for the patients most at risk, which is precisely what the AEQUITAS training is designed to do.

    References:

    Chang, C.T., Srivathsa N., Bou-Khalil, C., Swaminathan, A., Lunn, M.R., Mishra, K., Koyejo, S,. Daneshjou, R. (2025). Evaluating anti-LGBTQIA+ medical bias in large language models. PLOS Digit Health 4(9): e0001001. https://doi.org/10.1371/journal.pdig.0001001

    Cirillo, D., Catuara-Solarz, S., Morey, C., Guney, E., Subirats, L., Mellino, S., Gigante, A.A., Valencia, A., Rementeria, M.J., Chadha, A.S., & Mavridis, N. (2020). Sex and gender differences and biases in artificial intelligence for biomedicine and healthcare. npj Digit. Med. 3(81). https://doi.org/10.1038/s41746-020-0288-5

    Cross, J. L., Choma, M. A., & Onofrey, J. A. (2024). Bias in medical AI: Implications for clinical decision-making. PLOS Digital Health, 3(11), e0000651. https://doi.org/10.1371/journal.pdig.0000651

    Funding research on women’s health. (2024). Nature Reviews Bioengineering, 2, 797–798. https://doi.org/10.1038/s44222-024-00253-7

    Hafner, F.S., Valdivia, A., Rocher, L. 2025. Gender Trouble in Language Models: An Empirical Audit Guided by Gender Performativity Theory. FAccT ’25: Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, 1677–1695. https://doi.org/10.1145/3715275.3732112

    Karter, A. J., Parker, M. M., Moffet, H. H., & Gilliam, L. K. (2023). Racial and Ethnic Differences in the Association Between Mean Glucose and Hemoglobin A1c. Diabetes Technology & Therapeutics, 25(10), 697–704. https://doi.org/10.1089/dia.2023.0153

    Pennebaker, J. W., Mehl, M. R., & Niederhoffer, K. G. (2003). Psychological Aspects of Natural Language Use: Our Words, Our Selves. Annual Review of Psychology, 54(1), 547–577. https://doi.org/10.1146/annurev.psych.54.101601.145041

    Sjoding, M. W., Dickson, R. P., Iwashyna, T. J., Gay, S. E., & Valley, T. S. (2020). Racial Bias in Pulse Oximetry Measurement. New England Journal of Medicine, 383(25), 2477–2478. https://doi.org/10.1056/NEJMc2029240

    Smith, D.T., Mouzon, D.M., & Elliott, M. (2018). Reviewing the assumptions about men’s mental health: An exploration of the gender binary. American Journal of Men’s Health, 12(1), 78–89. https://doi.org/10.1177/1557988316630953

    Bowser, R. (2001). Racial bias in medical treatment. Dick. L. Rev., 105(3), 365.

    Williams, D. R., & Wyatt, R. (2015). Racial Bias in Health Care and Health: Challenges and Opportunities. JAMA, 314(6), 555. https://doi.org/10.1001/jama.2015.9260

  • The Hidden Ways AI Can Discriminate

    The Hidden Ways AI Can Discriminate

    When we talk about bias in AI, we are referring to “computer systems that systematically and unfairly discriminate against certain individuals or groups in favour of others” (Friedman & Nissenbaum, 1996). This is not about occasional errors, but about consistent, predictable patterns of unfair outcomes. Furthermore, biases in AI systems are complex, as they can enter the system at almost any stage, from the data it’s trained on to the way it’s ultimately used in a hospital setting. Recognising where bias originates, and how different types can reinforce each other, is the first step towards addressing it.

    Three broad roots of bias

    Researchers have identified three broad categories of bias that apply to all computer systems. The first is pre-existing bias, which comes from existing inequalities in society that get absorbed into the system, sometimes without anyone realising it (Friedman & Nissenbaum, 1996). The second is technical bias, which arises from the practical compromises made when translating complex human realities into clean computational form. And the third is emergent bias, which appears only after a system is deployed, as the world around it changes in ways the original design never anticipated.

    Bias in the AI pipeline

    These three categories give us a useful starting point. But for AI and machine learning specifically, researchers have identified seven more precise bias types that can affect these systems throughout their development and use (Suresh & Guttag, 2021):

    Historical bias reflects prejudices and stereotypes already present in training data, even if the data is technically accurate. A 1990 study found that after coronary bypass surgery, male patients received pain medication significantly more often than female patients, who were instead given sedatives more frequently (Calderone, 1990). An AI trained on data like this would learn to replicate that same pattern, perpetuating the very inequality it inherited.

    Representation bias happens when certain groups are underrepresented in training data. An AI trained mostly on data from one demographic will simply perform worse for everyone else. A well-known example is skin cancer detection tools that are significantly less accurate for patients with darker skin tones because the training datasets contained predominantly images from fair-skinned individuals (Guo et al., 2021).

    Measurement bias creeps in when the way something is measured isn’t equally accurate or fair across groups. One striking case involved an algorithm that used healthcare costs as a proxy for how sick someone actually was, and that ended up disadvantaging Black patients, who, facing disproportionate levels of poverty, tended to spend less on healthcare than equally sick white patients (Obermeyer et al., 2019).

    Aggregation bias occurs when a single, one-size-fits-all model is applied to genuinely diverse populations, ignoring the fact that the same data point can mean very different things depending on a person’s background.

    Learning bias emerges from technical decisions made while building the model itself, choices that can amplify disparities already present in the data, sometimes without the developers being aware. Research has shown, for instance, that differential privacy, a technique meant to protect patient confidentiality, can end up reducing a model’s accuracy for underrepresented groups even further (Bagdasaryan & Shmatikov, 2019).

    Evaluation bias happens when the datasets used to test and benchmark the model don’t reflect the real population it will actually serve, meaning a system can look accurate on paper while quietly failing specific groups of patients in practice.

    Finally, deployment bias arises when a tool is used in a context very different from the one it was designed and tested for, undermining its reliability in ways that are easy to miss.

    Why this matters for patients

    Understanding where bias comes from is the first step to catching it, and to making sure the AI tools shaping modern medicine work fairly for everyone, not just the patients who happen to be well represented in the data.

    This is exactly the kind of practical, structured understanding the AEQUITAS training equips healthcare professionals with, helping them recognise these patterns before they translate into real harm for real patients.

    References:

    Bagdasaryan, E., & Shmatikov, V. (2019). Differential Privacy Has Disparate Impact on Model Accuracy (arXiv:1905.12101). arXiv. https://doi.org/10.48550/arXiv.1905.12101

    Calderone, K. L. (1990). The influence of gender on the frequency of pain and sedative medication administered to postoperative patients. Sex Roles, 23(11), 713–725. https://doi.org/10.1007/BF00289259

    Friedman, B., & Nissenbaum, H. (1996). Bias in computer systems. ACM Transactions on Information Systems, 14(3), 330–347. https://doi.org/10.1145/230538.230561

    Guo, L. N., Lee, M. S., Kassamali, B., Mita, C., & Nambudiri, V. E. (2021). Bias in, bias out: Underreporting and underrepresentation of diverse skin types in machine learning research for skin cancer detection-A scoping review. Journal of the American Academy of Dermatology, S0190-9622(21)02086-7. https://doi.org/10.1016/j.jaad.2021.06.884

    Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447–453. https://doi.org/10.1126/science.aax2342

    Suresh, H., & Guttag, J. (2021). A Framework for Understanding Sources of Harm throughout the Machine Learning Life Cycle. Equity and Access in Algorithms, Mechanisms, and Optimization, EAAMO ’21, 1–9. https://doi.org/10.1145/3465416.3483305

  • AI Meets Healthcare: Understanding the Promise and the Risks

    AI Meets Healthcare: Understanding the Promise and the Risks

    Artificial intelligence is no longer a distant, futuristic concept in medicine. It’s already here, and is becoming an everyday part of clinical practice, supporting diagnosis, clinical decision-making, treatment planning, and patient risk prediction. In fact, according to the World Health Organization, nearly three quarters of EU countries are now using AI-assisted diagnostics in some form, and that number is only expected to grow (Alowais et al., 2023; World Health Organization, 2026).

    As the use of AI promises to further revolutionize the sector, what should we know before trusting it with our health?

    How AI works

    At its core, biomedical AI relies on three main approaches. Machine learning allows systems to learn patterns from data and make predictions without being explicitly programmed for every scenario. Deep learning, a more advanced form of machine learning, uses layered neural networks (loosely inspired by the human brain) to process complex information like medical images or clinical notes. And natural language processing (NLP) enables computers to understand and interpret human language, such as a doctor’s written notes or a patient’s description of their symptoms.

    Real examples already in use

    In cardiovascular care, AI tools are helping detect arrhythmias from ECG readings and estimating a patient’s individual risk of heart disease. In diabetes management, AI-powered retinal scans can catch signs of diabetic retinopathy earlier than ever before. And in mental health, natural language processing tools are being explored to help identify signs of depression in the way people write or speak, and to support patients between therapy sessions.

    So what’s the catch?

    Research has shown that AI applications can lead to misinterpretation and the promotion of biases and discrimination (Varona & Suarez, 2022). In essence, AI systems are only as good as the data they learn from. If that data reflects existing inequalities, for example, if it comes mostly from male patients, or predominantly from one ethnic group, the AI will inherit those biases. It will perform well for the people it “knows,” and poorly for everyone else.

    Several documented cases have shown AI tools trained on biased datasets producing less accurate results for underrepresented patients, sometimes with real consequences for diagnosis and treatment.

    What AEQUITAS is doing about it

    This is exactly the challenge our AEQUITAS project, funded by the European Union, was designed to address. By developing training resources, a database documenting real cases of bias in biomedical AI, and a practical regulatory model for hospitals and CSOs, AEQUITAS is helping healthcare professionals and institutions across Europe use these powerful tools more safely, and more fairly, for every patient.

    Understanding how AI works, and where it can go wrong, is the first step toward making sure it works for everyone. 

    References:

    Alowais, S.A., Alghamdi, S.S., Alsuhebany, N. et al. Revolutionizing healthcare: the role of artificial intelligence in clinical practice. BMC Med Educ 23, 689 (2023). https://doi.org/10.1186/s12909-023-04698-z

    Varona, D., & Suárez, J. L. (2022). Discrimination, Bias, Fairness, and Trustworthy AI. Applied Sciences, 12(12), 5826. https://doi.org/10.3390/app12125826

    World Health Organization. 2026. “Artificial intelligence is reshaping health systems: state of readiness across the European Union.” World Health Organization.