How Scientific American Went from Fighting Pseudoscience to Platforming It
The Strange Story of AI Tongue Diagnosis
I subscribed to Scientific American for most of my adult life. What attracted me to it was cool stories about topics that I would otherwise never encounter in my studies - new discoveries in space, paleontology, geology, and other topics that a veterinarian-turned-sport medicine professor would never encounter in my own research.
So when Scientific American began drifting into increasingly politicized territory a few years back, I eventually let my subscription lapse. This was not just my own observation, others had written about it also.
It wasn’t about my own politics (which, like most people’s, vary topic by topic). Much as I don’t go to POLITICO to learn about black holes, or Sports Illustrated for movie reviews, I didn’t subscribe to Scientific American to be told what conclusions I should reach on political issues. I subscribed because I wanted to learn about cosmology, neuroscience, physics, and other scientific topics that lay far outside my own field. When I wanted political analysis (or advocacy), I went elsewhere. But every now and then I check back in. Old habits die hard.
And recently, something caught my eye.
A “New” Scientific American Article on Tongue Diagnosis
I stumbled upon a hard copy of the October 2025 issue in a waiting room and began reading. Inside it, I found a piece that highlighted a study claiming that tongue photographs, analyzed by machine learning, could diagnose diseases.
My immediate reaction was: This can’t be serious. See below:
Tongue-reading has a long history in Traditional Chinese Medicine (TCM), and while it can be culturally meaningful, its scientific basis has never passed basic clinical scrutiny. So I clicked through to the original research article.
Within seconds, three red flags were waving so loudly they practically slapped me in the face.
The Claims Were Wildly Implausible
The paper, and news stories surrounding it, claimed that simple tongue photos could diagnose:
Diabetes
Mycotic infections
Asthma
Anemia
COVID-19
Even “resistant pylori infection” and “circulatory problems”
All based on tongue color categories like yellow, white, green, blue, and red.
As soon as I saw this, I knew the underlying methodology had to be flimsy. And unfortunately, it was worse than flimsy — it was almost comically so.
The “AI” Wasn’t Diagnosing Disease at All
When I opened the methods section, it became clear the model wasn’t diagnosing diseases.
Instead, the model was classifying colors.
Literally just colors. That’s right, AI was examining an image (not necessarily of a tongue) and saying “red” or “blue” or “yellow.” In other words, a skill that a 3-year old can reliably perform.
After that, it took images of 60 tongues and determined their color. I emphasize, that’s all it did, classify their tongue as one of seven colors.
Then, and this is the crucial part, the authors manually created a lookup table based on TCM that said things like:
Yellow tongue → diabetes
Green tongue → fungal infection
Blue tongue → asthma
Red tongue → COVID-19 or appendicitis or stroke (take your pick!)
In other words, the “AI model” did not learn to diagnose disease.
It learned to detect colors, and the authors manually attached Traditional Chinese Medicine disease associations to those colors. The leap from “yellow tongue” to “diabetes” wasn't learned by the AI, nor was it validated by the study. It was simply assumed.
Here is the chart they used:
And yet the paper presented the whole system as a “real-time disease diagnosis” tool with 96.6% accuracy, based on just 60 patients. There was no explanation of patient selection, no control matching, no confusion matrices, no clinical standards for diagnosis — nothing.
The algorithm wasn't evaluated on whether diabetes causes yellow tongues. It was evaluated on identifying whether yellow tongues are… yellow. That may sound flippant, but it is literally what the reported accuracy statistics are measuring — NOT accuracy of medical diagnosis. (And don’t even get me started on exploiting diagnostic accuracy statistics, see my previous work on that).
Fancy Math, Less Impressive Science
The most remarkable thing about the study is that all of the sophisticated mathematics distract from a much simpler reality.
The authors report support vector machines, random forests, XGBoost, k-means clustering, Jaccard indices, Fowlkes-Mallows indices, and an alphabet soup of machine-learning performance metrics.
See all of those formulae below? It may cause one to think that the analysis is so complex, it all must be meaningful.
Let’s add to that some images demonstrating the algorithms used, like naive Bayes classification:
This is machine learning. Yes, it all sounds, and looks, incredibly impressive… convincing even. But machine learning isn’t magic.
The algorithm was solving a trivial problem, identifying colors, while the much harder biological question went unexplored. None of these analyses answer the question most readers assume is being tested: can tongue color actually diagnose disease?
The study never asks whether:
most people with diabetes have yellow tongues, or
people without diabetes can also have yellow tongues, or
whether tongue color is even a reliable biomarker in the first place.
In effect, the AI gets credit for diagnosing diabetes without ever demonstrating that diabetes is associated with yellow tongues. The difficult biological question of whether tongue color actually predicts disease, is simply assumed to be true based on TCM.
All of the fancy mathematics, no matter how sophisticated, cannot rescue that fundamental problem. Identifying colors and diagnosing disease are very different tasks.
Clinical Utility (or Lack Thereof)
Ironically, the screenshots of the software reveal just how misleading the term “diagnosis” really is. Check out the Red Tongue diagnosis below:
A red tongue doesn’t lead to one diagnosis. Instead, it leads to a grab bag of possibilities.
Imagine it playing out in a doctor’s office. You sit down, stick out your tongue, and after a sophisticated AI analysis the physician returns with your results:
“We determined that your tongue is red. Unfortunately, that means you might have COVID-19, an acute stroke, appendicitis, H. pylori infection (a resistant one, nonetheless), or inflammation of the tongue.”
Naturally, you ask how the machine arrived at such wildly different possibilities… or the potential next step in the work-up. Perhaps you need a gastric endoscopy, an abdominal CT scan, and a brain angiography to narrow down the tongue-based differential list. Never mind that you had just spent the morning sucking on cherry cough drops (loaded with Red 40) to soothe the sore throat from your common cold.
That may sound ridiculous, but that’s remarkably close to what the system actually does. The AI isn’t deciding among those possibilities or even using clinical presentation to rank probabilities. It’s just telling you what color your tongue is, and then spitting out one of seven possibilities of what humans have programmed it to say.
That’s a bit like an AI recognizing that a rash is red and then proclaiming it could be poison ivy, lupus, cellulitis, a drug reaction, a snake bite, ringworm, shingles, eczema, PUPS, or a nasty rugburn. If it turns out to be any one of those, it is counted as a “correct” diagnosis.
But, It Goes Beyond Scientific American
To be fair, Scientific American wasn’t alone in covering this study. Popular Science, Forbes, India Times, and numerous other outlets enthusiastically covered the same story. The study’s Altmetric was >500, indicating it generated substantial attention.
My issue is that Scientific American presents the study largely on its own terms, provides only modest caveats, and fails to critically examine the extraordinary assumptions required to translate tongue color into disease diagnosis. The article does contain skepticism around clinical implementation, but very little skepticism around the underlying premise.
Weak studies become exaggerated by news outlets all the time. But I don’t subscribe to those publications, and I don’t expect them to hold themselves to the same standard in covering science. Scientific American is supposed to be different. That’s not just my opinion; they say so themselves.
Right below the article, they state they provide the “science world’s best writing and reporting” and that they cover “meaningful research and discovery.” Those are admirable aspirations. They also make pieces like this all the more disappointing.
It’s also worth noting that it’s not just Scientific American that sees itself that way. Media Bias/Fact Check, which evaluates news sources across the political spectrum, currently rates Scientific American as “High Credibility” (see report below).
Publications that have earned a reputation for exceptional rigor over generations should apply exceptional skepticism before presenting extraordinary claims to millions of readers. In my opinion, this study was so fundamentally flawed that it deserved far more scrutiny than it received in what is supposed to be one of the world's leading science magazines.
The lack of a critical lens is especially surprising, given some additional context. This was the magazine that once warned against legitimizing Traditional Chinese Medicine diagnoses and famously described acupuncture as “full of holes.” See the 2019 headline from the editors below:
The transformation from one that was previously willing to apply skepticism to TCM to one that treated this tongue-diagnosis study relatively uncritically is precisely why the article disappointed me.
And perhaps that’s the larger lesson. Scientific American isn’t uniquely flawed. It’s simply part of the same scientific storytelling ecosystem as everyone else. A provocative paper becomes a news story. Other outlets amplify it. Headlines become increasingly confident. Nuance disappears. Even Scientific American prominently featured the 96% accuracy claim in its subtitle.
I just expected better.
Why This Matters
For those unfamiliar with my writing, I am not interested in undermining science. Quite the opposite. And, I recognize that no study is perfect.
My concern is that public trust in science depends on our willingness to recognize when scientific communication falls short. Calling out weak evidence and poor reporting isn't anti-science. It's part of how science corrects itself and earns trust. We can, and should, do better.
Science journalism used to serve as a filter — messy ideas got distilled into clear, nuanced, evidence-based explanations for the public.
But when that filter stops working, or gets calibrated around “wow, that’s fun and interesting” at the expense of rigor, misinformation doesn’t need to be deliberately deceptive. It only needs to go unchecked.
A piece that lends credibility to a tongue-color diagnostic algorithm is exactly the kind of thing that erodes public trust in science journalism… and perhaps science.
At a time when misinformation is everywhere, science publications don’t just need to be accurate; they need to be judicious. They need to demonstrate that they still know the difference between evidence and appealing narrative, between research and wishful thinking.
The Takeaway
I didn’t go looking to criticize Scientific American. I just opened an issue for old times’ sake, because it was there.
But what I found (immediately) was a case study in how scientific standards slip when editorial filters weaken.
I want to read about science, the rigorous kind, and I know what it looks like.
I also know what it doesn’t look like. And this wasn’t it.
That doesn't mean Scientific American is worthless, or that every article it publishes should be viewed with suspicion. It means that no magazine article can substitute for evaluating the underlying evidence itself. Trusting science journalism doesn't mean turning off our critical faculties. It means using them.
.I hope you enjoyed this article. If you did, please consider subscribing and restacking. Also, check out my other Substack newsletters (all are free, and never have any spam or sales).
Academic Life: Essays, reflections, and advice all about academia and higher education.
Human Limits: Sports science, sports medicine, and exercise training insight.
Everyday Culture: Sometimes serious, sometimes zany. A look at parts of our culture that made me think, and inspired me to write.
Also, I’d love to hear your thoughts on this article! Please feel free to comment below.











Scientific American went from explaining science to preaching ideology while calling it “science communication.” They’ve redefined 'peer-reviewed' as 'progressive-approved' and 'facts' as 'facts that align with our values.'"
Great info. Reminds me of the ionic foot detox where you put your feet in water with an electrical device, the water turns brown, and it's supposedly drawing "toxins" out through your feet.
The brown color is just the metal electrodes corroding from electrolysis. It happens even without feet in the water. Your skin isn't a meaningful route for toxin elimination; that's your liver and kidneys. No credible research supports it. Classic detox marketing dressed up in wellness language. And I had a friend recently do it and was talking about how the water turned brown because all these toxins were leaving her body. I wanted to tell her the truth but...