Solutions
Low abundance feature recovery
The features that matter most are often the ones closest to the baseline. Thresholds that keep noise out also discard real low abundance signal, and the loss is invisible because a missing feature leaves no trace.
The problem
Low abundance feature recovery
Low abundance features sit near the noise floor, where the choice of intensity threshold decides what survives. Set it high and real compounds vanish. Set it low and the table fills with noise that has to be cleaned later. Because a real but weak feature that falls below threshold simply does not appear, its absence is silent, and it can be the exact biomarker candidate a study was designed to find.
Why it happens
Where conventional workflows produce it
Single sample detection has to distinguish signal from noise using one trace, so a conservative threshold is the safe default. That default is applied uniformly, which penalizes the weakest real features hardest. Noise is random and does not reproduce across injections, but a per sample picker cannot use that fact because it never looks across the cohort.
How Metablify addresses it
Evidence from the whole cohort
Metablify uses reproducibility as the discriminant. A weak feature that appears at consistent mass and retention across many injections is unlikely to be noise, even when its intensity in any single run is small. By amplifying that agreement, the platform recovers low abundance features that a fixed threshold would reject, while random fluctuations that do not reproduce stay out. The law of large numbers does the work that a single trace cannot.
What changes
What you see in the output
More real features near the baseline enter the table with defensible evidence, fewer true positives are lost to a blunt threshold, and studies that depend on subtle differences keep the signal they were designed to detect.
Honest limits
What this does not do
A feature that was never sampled above noise in any injection cannot be recovered, because there is nothing consistent to amplify. Recovery improves with more injections and with a cohort that actually contains the feature. It does not lower the instrument's fundamental limit of detection.
What matters
Where this makes a difference
Silent losses are the danger
A real feature below threshold simply does not appear, so its absence leaves no trace and can be the exact biomarker candidate a study was designed to find.
Reproducibility separates signal from noise
Noise is random and does not recur. A weak feature seen at consistent mass and time across injections is unlikely to be noise, which is how recovery rises without lowering a threshold.
More injections, more evidence
Because the discriminant is agreement across the cohort, larger and more replicated studies recover more low abundance features rather than fewer.
Questions
Common questions
Does recovering weak features add noise to my table?
The point of using cohort agreement is to separate the two. Random noise does not reproduce across injections, so it is not amplified. Weak but reproducible features are, which raises recovery without simply lowering a threshold.
How many samples do I need for this to help?
More injections give more evidence, so larger cohorts benefit most. Even modest replication helps, because the discriminant is reproducibility rather than raw intensity in a single run.
Sources
- Mahieu and Patti, feature reliability, Analytical Chemistry (2017) ↗How many features in untargeted data are real.
- Limit of detection, IUPAC recommendations ↗Definitions for detection near the baseline.
Keep reading
Related
Prove it on your own data
Send a limited set of your existing LC/MS data and see how many additional real mass features are recovered against your current output.
Send replicated data and see how many reproducible low abundance features are recovered against your current output.