Eight years on one narrow question.
ContextRx is an independent research project by Kevin Kimathi exploring the use of NLP and machine learning to detect contextual word errors in pharmaceutical documents and drug labels — real words, correctly spelled, wrong for the clinical context they appear in.
It has been in development for eight years across five architectural iterations, from rule and n-gram baselines through statistical language models to learned contextual representations fitted on a corpus of 34,201 FDA-approved label documents.
It currently operates as a production system scanning FDA-approved drug labels and identifying errors that pass through existing pharmaceutical quality systems, including regulatory review. That result — that this error class survives every layer of review and reaches the published label — is the finding the project is built around.
Why this class of error
Spellcheckers were designed for the office document, and general-purpose grammar tools for everyday prose. Neither is built to know that "patents" in a dosing instruction, or "temperate" in a storage condition, is a clinically wrong word. The error is invisible precisely because it is spelled correctly.
Independence and scope
This is personal research conducted independently, on personal time and infrastructure. It is not a commercial product, it is not for sale, and it is not affiliated with, sponsored by, or endorsed by any employer or client. Findings are research observations only and are not regulatory advice or a determination about any product, manufacturer, or submission.
Peer review, replication attempts, methodology critique, and collaboration are all genuinely welcome.