Articles

Metabolite Annotation and Databases for MS Imaging

The hardest part of spatial metabolomics is not acquiring the data but knowing what each peak is. And one of the standard axes of evidence is permanently unavailable.
Written byTrevor J Henderson
A bioinformatician at two monitors, one showing a long list of candidate metabolite annotations and the other a single tissue map.

One ion image, many candidate structures. Which of them you may legitimately name depends on evidence imaging cannot always supply.

Flow (2026)

Doing metabolite annotation in MS imaging well begins with accepting a constraint that is rarely stated directly: the community framework for identification confidence was built around chromatography, and imaging has none. That is not a gap awaiting better instruments. It is a structural feature of the technique, and understanding it changes what you should claim and how you should spend your effort.


Key Takeaways

  • The community confidence framework runs from level 5, exact mass only, to level 1, a structure confirmed against an authentic standard.
  • Level 1 requires chromatographic retention time. Imaging has no chromatography, so level 1 is structurally unattainable rather than merely difficult.
  • That makes ion mobility disproportionately valuable in imaging — collision cross-section is the only available replacement for the missing axis.
  • Database choice largely determines how many candidate structures each annotation carries, so a curated database beats a comprehensive one.
  • In-source fragments masquerade as distinct metabolites, and spatial co-localisation is the practical way to catch them.

Why Annotation Is Hard in MS Imaging

Three difficulties compound, and only the first is shared with conventional metabolomics.

  • Mass does not determine identity. Isomers share a molecular formula exactly, so no improvement in mass accuracy separates them. This is true of all mass spectrometry.
  • There is no separation before ionisation. Everything present at a pixel is ionised together, which both crowds the spectrum and removes the separation dimension that conventional workflows rely on, as set out in Mass Spectrometry Imaging: Principles, Techniques, and Applications.
  • Fragmentation is expensive. On-tissue tandem MS is possible but costs acquisition time at every pixel where it is applied, so it is typically reserved for selected targets rather than applied across an image.

The consequence is that most imaging annotations rest on MS1 evidence alone, which is a weaker position than a routine LC-MS experiment occupies. The scale of the resulting gap is documented: as discussed in Spatial Metabolomics and Lipidomics by Mass Spectrometry Imaging, the developers of the field’s standard annotation engine describe the unassignable majority of imaging data as dark matter.

One terminological warning before going further. In this literature, the abbreviation MSI means both mass spectrometry imaging and the Metabolomics Standards Initiative, and both appear in discussions of annotation confidence. This article spells out the standards body in full throughout to avoid the collision.

Working in analytical science?

Register for a FREE Separation Science account to subscribe to the Separation Science Newsletter.

Subscribe for free

Why Can’t MS Imaging Reach Level 1?

Because level 1 was defined to require something imaging does not produce. This is worth working through carefully, since it determines what you can honestly claim.

The Metabolomics Standards Initiative originally proposed a four-tier system for communicating identification confidence, which was subsequently refined by Schymanski and colleagues into a five-level framework tailored to the capabilities of high-resolution mass spectrometry. The levels run from level 5, exact mass alone, to level 1, a confirmed structure supported by a reference standard, incorporating molecular formula, spectral library matches, and structural inference along the way.

The top of that scale has a specific requirement. Level 1 is reserved for compounds confirmed by direct comparison with an authentic standard measured under identical analytical conditions, and the standards initiative definition requires a minimum of two independent and orthogonal pieces of data from that standard. In practice, for LC-MS work, those axes are accurate mass, fragmentation, and chromatographic retention time.

Level

What It Represents

Evidence Required

Available in Imaging?

5

Exact mass only

Accurate mass measurement

Yes. This is the baseline for every imaging feature

4

Unequivocal molecular formula

Accurate mass plus isotope pattern

Yes, with high mass resolving power

3

Tentative candidate or compound class

Formula plus partial structural evidence

Yes, with on-tissue fragmentation or class-specific ions

2

Probable structure by library or diagnostic evidence

Spectral library match or strong structural inference

Partly. Library matching is harder without a separation dimension

1

Confirmed structure

Authentic standard under identical conditions, conventionally including retention time

No. There is no chromatography, so no retention time to match

Table 1. The five-level confidence framework mapped against what mass spectrometry imaging can supply. The framework is as published; the availability column is our assessment of what imaging can and cannot provide.


This Is a Structural Limit, Not a Maturity Problem

Coverage of imaging annotation often implies the field will catch up with LC-MS metabolomics as instruments improve. On this particular axis it will not, because the limitation is not sensitivity or resolving power but the absence of a separation stage. A technique with no chromatography cannot produce a retention time to match against a standard, however good its mass analyser.

Two useful conclusions follow. First, report honestly: an imaging annotation is normally a level 4 or level 3 claim, and describing it as an identification overstates it. Second, spend effort where it can actually move you — which means the axes imaging can still add, rather than the one it cannot.

Databases and Reference Resources

Database choice is the most consequential decision in an annotation workflow and the least deliberated. It sets both what can be found and how many candidates each finding carries.

Resource

Scope

How It Behaves in Imaging Annotation

HMDB, Human Metabolome Database

Broad coverage of human metabolites, including many not expected in a given tissue

Comprehensive, and therefore returns large candidate lists per annotated ion

LIPID MAPS

Lipid classification, nomenclature, and structures

Essential for lipid work, and supplies the notation that encodes structural evidence level

Expert-curated core databases

Deliberately restricted to metabolites plausibly present and detectable

Fewer candidates per ion and better precision; the METASPACE developers report improved results over general databases

Pathway and compound registries

Biochemical context and cross-references

Useful for interpretation rather than assignment

Specialised community databases

Domain-specific compounds, for example microbial specialised metabolites

Substantially improves annotation in contexts general databases cover poorly

Table 2. Reference resources for imaging annotation. The trade-off in the third column is the practical point: comprehensiveness and precision pull against each other.

The counterintuitive guidance is to prefer the smaller database. Because imaging annotation works from MS1 evidence, every additional structure in the database that shares a molecular formula with a real feature becomes another candidate you cannot exclude. A comprehensive database therefore inflates ambiguity without adding information. The METASPACE machine learning work reports exactly this: introducing an expert-curated core database improved results relative to general databases, and the model itself was trained and evaluated across 1,710 datasets from 159 researchers at 47 laboratories, spanning animal and plant contexts.

Continue reading below…
Infographics3D visualization of protein structures in top-down proteomics
Top-Down vs Bottom-Up Proteomics
Explore the essentials of Top-Down Proteomics in battery materials quality control. Download our insightful infographic today!
Read More

For lipids specifically, the LIPID MAPS classification and shorthand notation standard does double duty: it is a structural reference, and it provides the notation that declares how much structure you have actually established, which is developed in Spatial Lipidomics in Tissue.

How Does FDR-Controlled Annotation Work?

By asking how often the method would produce an annotation that cannot be real, and using that rate to calibrate confidence in the ones that might be.

The approach implemented for imaging generates implausible decoy ions, scores real candidate annotations against them, and reports annotations at a stated false discovery rate by that ranking. Scoring combines several measures, including agreement between measured and theoretical isotope patterns and the spatial coherence of the ion image, on the reasoning that a real molecular distribution is spatially structured while noise generally is not.

Three things follow that are worth understanding before quoting an FDR figure.

  1. The FDR is conditional on the database. It expresses the rate of false formula assignments given the search space you chose. Changing the database changes the figure, so an FDR without a stated database is uninterpretable.
  2. It controls formula assignment, not structure. An annotation passing a 10 percent FDR threshold is a formula reported at that confidence, still potentially corresponding to many isomeric structures.
  3. Spatial coherence is doing real work. This is a genuine advantage imaging has over infusion experiments: the image itself is evidence, because a spatially random distribution is unlikely to be a real metabolite localisation.

That third point deserves emphasis because it partly offsets the missing retention time. Imaging gains an evidence type that conventional metabolomics does not have, namely the spatial structure of the signal. It is not equivalent to a separation dimension, since it speaks to whether a signal is real rather than to what it is, but it is not nothing.

Isomers and In-Source Fragments

Two categories of ambiguity behave differently and require different responses. Isomers are a structural problem; in-source fragments are an artefact problem, and the second is more often overlooked.

In-source fragmentation occurs when a molecule breaks apart during desorption and ionisation rather than in a collision cell. The resulting fragment appears in the MS1 spectrum as though it were an independent species, and it will be annotated as one if its mass matches a database entry. In imaging, this is particularly troublesome because there is no chromatographic separation to reveal that the fragment and its parent share an origin.

Continue reading below…
Learning HubsComplex peptide-like molecular structures representing large drug metabolites analyzed by LC-MS/MS
Build Confidence in Metabolite Identification
Discover time- and money-saving solutions that generate reliable data without compromising precision or efficiency.
Read More

Three practical defences, in order of usefulness.

  • Check spatial co-localisation. A fragment will map exactly onto its parent, because it is generated from it at the point of desorption. Perfect co-localisation between a smaller and a larger species related by a plausible neutral loss is a strong indication of in-source fragmentation rather than two co-regulated metabolites.
  • Look for characteristic neutral losses. Losses of water, phosphate, or head groups relate fragments to parents predictably within a class, so they can be anticipated.
  • Moderate the ionisation energy. Reducing laser fluence lowers in-source fragmentation, at some cost in signal, which is a trade worth testing during method development.

The first of those is the one to internalise, because it inverts an instinct. Two ion images that overlay perfectly look like a strong biological result and are frequently an artefact of one molecule being counted twice. Suspiciously exact co-localisation warrants checking rather than celebration.

How Do You Improve Confidence Without Retention Time?

By adding the axes that remain available. Given the structural argument above, this is where method development effort actually pays, and one option matters more than the others.

Approach

What It Adds

Practical Cost

High mass resolving power

Confident molecular formula assignment, reaching level 4

Acquisition speed, which matters given pixel counts

Ion mobility and collision cross-section

A genuinely orthogonal structural descriptor, and the closest available substitute for retention time

Modest. Separation is achieved without appreciably slowing the raster

On-tissue tandem MS

Structural evidence toward level 3 or 2

Acquisition time, so usually applied to selected targets or regions

Curated database restriction

Fewer candidates per annotation, improving precision

Risk of excluding a genuinely present but unlisted compound

Spatial coherence filtering

Confidence that a feature is real rather than noise

Little, and it is imaging-specific evidence worth exploiting

Orthogonal validation by extraction

Can reach level 1 for selected compounds, since the extract can be chromatographed

Destroys spatial information for that sample, so it is a separate experiment

Table 3. Routes to higher annotation confidence in imaging. The second and last rows are the two most strategically important.

Ion mobility deserves the emphasis for a specific reason. Conventional metabolomics has accurate mass, fragmentation, and retention time as orthogonal axes, with collision cross-section increasingly added as a further one. Imaging permanently lacks retention time, so cross-section is not an incremental refinement here but the replacement for a missing dimension. That is why it carries more weight in imaging than in LC-MS work, where retention time already does that job. A spatial metabolomics workflow using cyclic ion mobility with predicted collision cross-sections reported accuracy better than 0.4 percent relative to database values in multipass experiments, which was sufficient to improve filtering thresholds and resolve assignments that mass alone left ambiguous.

The final row of Table 3 is the honest route to a definitive identification, and it is worth stating explicitly: extract the region of interest and run a conventional separation-based experiment on it. That recovers retention time and can reach level 1, at the cost of the spatial information for that sample. Discovery by imaging followed by confirmation by extraction is not a workaround, but the appropriate division of labour, and it is the workflow pattern recommended in Untargeted Spatial Metabolomics Workflows.

Software, processing choices, and computational approaches to all of this are covered in Analyzing Mass Spectrometry Imaging Data: Processing, Statistics, and Multimodal Integration. For where annotation sits in the wider spatial landscape, see Spatial Analysis in Analytical Science: Mass Spectrometry Imaging and Spatial Omics.

This article was produced under Separation Science's AI Editorial Guidelines.

Frequently Asked Questions (FAQs)

  • How do you annotate metabolites in MS imaging?

    By matching measured accurate masses against a database of molecular formulas, scored and reported at a controlled false discovery rate using implausible decoy ions as a reference. Scoring typically combines isotope pattern agreement with the spatial coherence of the ion image. Because the evidence is normally MS1 only, an annotation is a formula with a candidate list rather than a confirmed structure.

  • What is METASPACE?

    The community annotation engine for imaging mass spectrometry. Users upload imaging datasets and receive metabolite and lipid annotations at a controlled false discovery rate, scored against a chosen database of molecular formulas. Its developers have since added a machine learning approach trained across 1,710 datasets from 159 researchers at 47 laboratories, along with an expert-curated core database that improved precision over general databases.

  • How do you control false discovery in MS imaging?

    By generating implausible decoy ions, scoring real candidate annotations against them, and reporting annotations at a stated false discovery rate based on that ranking. Two caveats matter: the rate is conditional on the database searched, so it is uninterpretable without one stated, and it controls molecular formula assignment rather than structure, so isomeric ambiguity remains.

  • Can MS imaging identify metabolites definitively?

    Not to the highest community confidence level in a single experiment. Level 1 identification requires comparison with an authentic standard under identical conditions, conventionally including chromatographic retention time, and imaging has no chromatography. The honest route is discovery by imaging followed by extraction of the region of interest and a conventional separation-based experiment for confirmation.

Add Separation Science as a preferred source on Google

Add Separation Science as a preferred Google source to see more of our trusted coverage

Meet the Author(s):

  • Trevor Henderson

    Trevor Henderson, PhD, is a veteran Content Innovation Director and scientific strategist at LabX Media Group. With a career spanning three decades, Trevor is a recognized expert in scientific writing, creative content creation, and technical editing.

    His academic pedigree in human biology, physical anthropology, and community health provides him with a rigorous analytical framework, which he applies to developing industry-leading content for scientists and lab technicians. Since 2013, Trevor has led content innovation initiatives that drive engagement within the laboratory technology sector.

    View Full Profile

Here are some related topics that may interest you:

Related Content