AI skin cancer detection tools are getting better – but only for people with light skin
These apps and software programs could provide access to lifesaving screening – if researchers can find a way to remove the baked-in bias.

Imagine you’re getting out of the shower one morning and you notice a mole on your thigh that you’ve never seen before. It’s reddish brown, bumpy and surprisingly large. Is it a benign mole, or is it melanoma?
A slew of new artificial intelligence tools claim they can help you figure it out. Some are smartphone apps that anyone can download to scan their skin at home, while others are software programs designed to be used directly by clinicians in a doctor’s office.
As a computer engineer studying how tools like these perform in real-world clinical settings, I know that finding a way to accurately use AI in dermatology would be immensely valuable to patients around the world. It could offer broad access to medical expertise, providing lifesaving screenings to remote areas or underresourced communities where dermatologists are scarce.
But at the moment, these tools have a crucial shortcoming that researchers will have to resolve: They are increasingly accurate for people with light skin – but they have a massive blind spot when it comes to analyzing darker skin.
Skin-deep accuracy
The central myth of AI is that it functions objectively. In reality, an AI model is simply a pattern-matching engine. It learns to associate certain visual features with certain diseases.
But it can easily be thrown off by the background color of a person’s skin. In other words, the AI model doesn’t learn to look at the lesion itself. Instead, it picks up on the color of the surrounding skin as a clue. This means that the model’s ability to make accurate predictions essentially degrades to guesses based on skin color.
My colleagues and I found that the ability of these tools to accurately diagnose skin conditions dropped significantly when we simply darkened the surrounding skin on patients’ images. We trained an AI model on photographs of known skin conditions in light-skinned patients, then digitally manipulated the images to resemble darker skin tones. The clinical condition in the photo had not changed, but the AI’s ability to recognize it deteriorated sharply.
For example, consider a condition such as atopic dermatitis – a chronic, itchy and inflammatory skin disease. It causes a skin discoloration that appears pink on light skin but gray or violet on darker skin. Our research and that of other groups shows that AI models would reliably classify the pink marks but might not identify the darker colors as signs of atopic dermatitis.
This AI classification blind spot means that patients with darker skin would receive measurably worse care. The disparity has critical consequences because skin cancers such as melanoma are visually harder to spot on pigmented skin. Patients of color are already more likely to be diagnosed at a more advanced stage, which leads to significantly lower survival rates. A diagnostic tool that works better for people with lighter skin only widens this gap.
The skin tone gap
This bias extends beyond tools used by clinicians to more widely used AI chatbots, such as ChatGPT or Claude. The stakes can become higher when people turn to these tools for medical answers without a clinician to double-check the output, making accuracy across all skin tones a matter of patient safety.
In a 2024 study, we presented OpenAI’s model, GPT-4, with an image of a completely benign mole. When we digitally darkened the skin around the mole while keeping the mole itself exactly the same, GPT-4 classified the spot as malignant melanoma. Because the skin color is the more prominent feature, the AI became so focused on the dark pigment of the skin that it ignored the standard medical rules used to identify cancer, such as checking whether the mole’s borders are irregular.
If a person uses these tools at home, a harmless dark spot might trigger unnecessary panic, while a life-threatening cancer on dark skin could be overlooked.
Fixing how AI is trained
Why does this bias exist? The answer lies in the images used to train AI models.
Researchers build these AI models by feeding them hundreds of thousands of images pulled from public online libraries of medical photos shared by universities and hospitals.
Historically, these medical databases – as well as dermatology textbooks – have been dominated by images of lighter skin tones. Darker skin images are relatively rare, in part because clinical norms were developed primarily around white patients. If an AI model is never taught what melanoma looks like on dark skin, it simply won’t know how to find it.
To achieve the same accuracy on darker skin as on lighter skin, these models need a more diverse set of images for training. However, while there are millions of photos of light-skinned patients already available in historical databases, gathering a massive new database of real photos from patients of color raises challenging ethical and patient privacy issues.
Generative AI offers limited help
One way around this privacy hurdle may be to use generative AI – the same technology powering chatbots, which can also be used to make deepfake videos – to artificially generate thousands of synthetic medical images.
Using just text prompts, researchers can create realistic, high-quality images of conditions such as melanoma on darker skin tones. My colleagues and I showed that an AI trained entirely on these synthetic images can learn to correctly categorize data using those images just as well as one trained on real ones.
However, this approach carries a hidden risk: Generative AI models can create high-quality images, but they may not map cleanly onto the characteristics of skin conditions that patients actually experience. This would be like using an inaccurate map to teach someone how to navigate.
If researchers train a diagnostic AI tool on these flawed synthetic images, the data might look perfectly diverse on a spreadsheet, but the tool will remain functionally blind to how these diseases actually appear on real patients of color.
At the moment, there is no shortcut around this problem. The only viable path forward is to build more inclusive and representative collections of images, particularly from people with darker skin tones.
Today, the medical AI field is at a crossroads. While some of these AI skin-scanning tools are already making their way into clinics and app stores for use in the U.S. and around the world, researchers and regulators are already pushing for stricter testing across all skin tones before these tools are widely deployed.
Ultimately, eliminating color-based bias in AI isn’t just about fairness, but rather the absolute baseline required to ensure these tools actually work for the people who need them most.
Mohamed Akrout does not work for, consult, own shares in or receive funding from any company or organization that would benefit from this article, and has disclosed no relevant affiliations beyond their academic appointment.
Read These Next
Stars and Stripes shows how government-funded journalism can work – if it’s free to report the truth
The military newspaper earned credibility by informing troops about the good, the bad and everything…
From the Freedom Caucus to the democratic socialists, small groups inside parties can wreak havoc wi
What role might DSA-affiliated candidates play if they win election to Congress in the midterms and…
Satellite images show how penguins’ diets are changing in response to Antarctic climate change
When penguin colonies are surrounded by sea ice, the penguins tend to eat more fish, rather than krill.…




