Liver Disease and Diagnostics · eu
Blood Tests and Elastography Each Have Their Strengths: Large Prospective Study Redraws the Noninvasive Diagnostic Landscape for Fatty Liver Disease
Using liver biopsy as a common reference, the LITMUS study compared multiple blood- and imaging-based tools. The results showed that serum markers performed better in identifying high-risk steatohepatitis, while elastography led in detecting advanced fibrosis and cirrhosis—but no single test could address every clinical question.
The real challenge of fatty liver disease is not merely the accumulation of fat in the liver, but identifying which patients have already developed inflammation, liver-cell injury, and fibrosis and therefore require more intensive follow-up or treatment. Although liver biopsy is the current reference standard, it is invasive and is also affected by sampling location and differences in pathological interpretation. The LITMUS Imaging Study, published in *Nature Medicine*, suggests that replacing biopsy may not depend on a single all-purpose test; instead, different tools may need to be selected according to the diagnostic objective.
This prospective, multicenter observational study enrolled 553 participants. The primary analysis included 357 patients with metabolic dysfunction-associated steatotic liver disease (MASLD) who had both central pathology review and imaging data obtained within a close time interval. The study compared magnetic resonance elastography (MRE), vibration-controlled transient elastography (VCTE), other magnetic resonance techniques, multiple serum biomarkers, and scores combining liver stiffness with clinical data. Each result was then compared with liver biopsy assessments of steatohepatitis and fibrosis stage.
When the question was how to identify metabolic dysfunction-associated steatohepatitis (MASH), or “high-risk MASH” with at least stage 2 fibrosis, the serum marker NIS2+ performed best, with areas under the receiver operating characteristic curve (AUCs) of 0.83 and 0.82, respectively. The closer an AUC is to 1, the better the ability to distinguish between those with and without disease. However, neither result statistically significantly exceeded the study’s prespecified minimum acceptable standard. In other words, ranking first numerically does not mean the test is sufficient on its own to determine who should receive treatment.
For fibrosis staging, the advantage shifted to technologies that measure liver stiffness. MRE achieved an AUC of 0.91 for both advanced fibrosis and cirrhosis; Agile 3+ had an AUC of 0.84 for advanced fibrosis. For cirrhosis, VCTE liver stiffness, Agile 3+, and Agile 4 had AUCs of 0.87, 0.89, and 0.88, respectively, and all crossed the prespecified performance threshold. This division of strengths is not unexpected: signals of inflammation and liver-cell injury in the blood more closely reflect disease activity, while elastography directly captures the tissue stiffening associated with fibrosis.
The findings could influence how patients are enrolled in clinical trials and selected for treatment. When seeking patients with high-risk MASH, more readily scalable blood tests could first narrow the pool of candidates. If the priority is confirming advanced fibrosis or cirrhosis, MRE, VCTE, or composite scores incorporating liver stiffness may be more appropriate. Actual workflows must still balance equipment availability, cost, and the consequences of misclassification, rather than being determined solely by AUC rankings.
The evidence also has clear limitations. All participants had been referred to secondary or tertiary care and undergone biopsy, and the prevalence of disease was higher than in the general community; the results therefore cannot be directly extrapolated to screening performance in primary care. In addition, 94% of participants were white, and applicability to other populations remains to be validated. The number of samples available varied across biomarkers, complete head-to-head comparisons covered only a subset of participants, and some confidence intervals overlapped. Given that liver biopsy itself may be subject to sampling and interpretation errors, and that multiple authors had conflicts of interest involving imaging, diagnostic, and pharmaceutical companies, these tools still need validation in more diverse routine-care populations.