helmlit

Coverage and limitations

What is in this corpus, what is not, and what has not yet been checked. Read this before citing anything found here in a review.

Precision is about 86%. Measured on a random sample of 300 PubMed records, labelled by title and then by abstract where the title was not decisive: 258 were on topic, 42 were not (±4 points, 95%). Recall against a known bibliography is about 99%. So roughly one record in seven is off topic — treat results as a screening set, not a finished one.

The noise is not random. Most of it is papers that mention C. elegans once in an unrelated molecular biology abstract, or use Ascaris antigen as an allergen in an asthma model. Both are admitted on purpose by the model-organism and immunology rules. A separate substring fault — “helminth” also matching the fungus Helminthosporium and the fern Helminthostachys — accounted for 535 records and has been removed.

What is included

Records in scope0across all sources below
With open full text0the rest is title and abstract only
Entity mentions0genes, organisms, chemicals, diseases, resolved to identifiers
Citation links0JATS references and NIH iCite

How scope was defined

A versioned rules file combining MeSH subject headings (pinned to the 2026 tree), about 60 organism and drug terms matched in titles and abstracts, and an exclusion for plant-parasitic and entomopathogenic nematology. Tiers such as immunology and microbiome only admit a paper alongside a core helminth term, so “eosinophils” alone does not pull in the allergy literature. Validated at 98% agreement with the equivalent live PubMed query.

Known gaps

Paywalled full textAbout 194,000 papers have no open full text. Searching here reaches their titles and abstracts only.
CAB AbstractsNot licensed. It covers 90% of veterinary journals against about 70% reachable without it — the single largest gap, and veterinary parasitology is roughly half this field.
Non-PubMed records lack MeSH and text miningWHO IRIS, Europe PMC-only and OpenAlex-only records are searchable by title and abstract, but they carry no MeSH indexing and no text-mined entities, so they will not appear on entity pages. Scope for them came from the harvest query, not the scope rules.
Citations and entities load more slowlySearch is fast, but the citation graph and entity mentions on a paper page are read directly from the full corpus file over HTTP range requests, so those panels can take a few seconds to fill in.
Gene and protein identifiers are sparseThe entity layer is strong for organisms (5.0M mentions, 72,368 distinct taxa) and chemicals (1.9M mentions), and thin for genes and proteins: about 147,000 mentions carry a real UniProt accession. A further 3.2M carried a placeholder identifier rather than an accession and are excluded, since they collapsed unrelated proteins onto a single entity. Entity pages are therefore reliable for organisms and compounds, and partial for genes.
Ensembl IDs here are humanThe Ensembl gene identifiers in this corpus (about 193,000 mentions) are host genes from immunology papers, not parasite genes. There is no shared gene identifier between this literature index and WormBase ParaSite assemblies, so a gene-level join to genome data needs a UniProt-to-WormBase crosswalk that does not exist here yet.
African Index MedicusNo public API. African regional journals are under-represented, which matters for a disease burden concentrated in sub-Saharan Africa.
Chinese-language literatureCNKI and Wanfang require paid licences. Most Schistosoma japonicum research is published in Chinese and is largely absent.
Pre-2000 materialOver half the paywalled remainder predates 2000, much of it in journals that no longer exist. No modern route reaches it.

How search behaves

Queries run server-side against a search index holding titles and abstracts. Because that index stores the searchable text without keeping a copy of it, results show no matched-term extract — the ranking is unaffected, but you will not see the sentence a term appeared in until you open the paper.

BM25 over titles and abstracts, title weighted four times. No stemming, deliberately: gene symbols and species binomials must match exactly, so daf-2 is not split and Necator is not conflated with necatoriasis. Entity pages use text-mined identifiers instead of keywords, so a paper appears under an organism whatever name it used. There is no semantic or embedding-based search — a query that paraphrases rather than names its subject may miss.

Retracted papers

Retracted papers are kept and flagged, never removed. A retracted paper is still evidence of what was claimed and when. They are marked in results and carry a banner on their page.