Research
Digital Medicines Information at National Scale: Search Behaviour and System Performance of MediVerify Across 1.5 Million Queries
Overview Research area: Digital health informatics and medicines information systems — specifically the analysis of search behaviour on a national medicines information platform, with relevance to pha

- arXiv
- 2609.28946
- Published
- 2026-09-24
- Authors
- Praveen Charuka Athauda-Arachchi, Pandula Mahesh Athauda-Arachchi, Rohini Fernandopulle
AI summary
Overview
Research area: Digital health informatics and medicines information systems — specifically the analysis of search behaviour on a national medicines information platform, with relevance to pharmaceutical policy, AI/health data ethics, and health system performance in low- and middle-income countries (LMICs).
Technical level: Intermediate. The statistical concepts are straightforward, but understanding the paper requires familiarity with fuzzy string matching, ATC therapeutic classification, and platform performance metrics.
Scope (one sentence): A retrospective, one-year analysis of 1,497,304 anonymised medicine queries on Sri Lanka's MediVerify platform, characterising what the public searched for, how that compares with essential medicines policy, and how well the system performed.
What This Paper Is About
National digital health platforms accumulate large volumes of usage data, but medicine search behaviour at scale has been poorly described in LMICs, so it is unclear what populations actually look for and whether that demand aligns with official medicines policy. This paper asks three questions using a year of MediVerify queries: what medicines people search for, where public demand diverges from essential medicines lists, and whether a national platform can serve that demand reliably and cheaply. It is an observational study of logs rather than a clinical trial or a model-development paper.
Key Contributions
-
A large-scale description of medicine information-seeking behaviour in an LMIC. The authors characterise queries submitted to a national platform over its first year, a setting the abstract describes as previously under-characterised.
-
A demand-versus-policy comparison. By mapping queries onto approved medicines and therapeutic categories, the study identifies where frequently searched medicines diverge from essential medicines lists — framed as a way to surface mismatches between public demand and policy.
-
A technical and environmental performance assessment. The paper reports query resolution rates, median latency, and an estimate of energy consumption derived from processing times, positioning the platform's efficiency as a finding in its own right.
-
A demonstration of usage analytics as policy input. The abstract argues that routine search logs can be repurposed to strengthen pharmaceutical policy and digital health infrastructure in LMIC settings.
Main Findings
-
Therapeutic demand concentrated in a few categories. Vitamins and minerals (11.74%), antibiotics (10.57%), anti-diabetes medicines (7.02%), and antihypertensives (6.84%) dominated searches.
-
Demand was concentrated but not monolithic. The top 20 medicines accounted for 37.3% of all queries.
-
A large share of the formulary went untouched. Roughly 40%–50% of registered medicines were never queried during the period studied.
-
A small but meaningful residue of unmet need. Zero-result queries made up about 1% of the total, which the authors interpret as evidence of unmet information needs.
-
Public demand and essential medicines lists do not fully align. Frequently searched medicines diverged from essential medicines lists, which the abstract frames as a policy-relevant mismatch.
-
Strong technical performance. Median query latency was 8 ms, with more than 99% of queries resolved.
-
Low estimated environmental cost. Estimated energy consumption was approximately 0.12 kWh per million queries.
-
Overall claim: a national digital health platform can achieve high utilisation at low computational and environmental cost, and its usage analytics can inform pharmaceutical policy.
Methodology in Plain English
The researchers looked backwards at one year of anonymised search queries submitted to MediVerify between July 2024 and 2025 — 1,497,304 queries in total — and treated them as a record of what the public wanted to know about medicines.
To make raw search text analysable, they cleaned and normalised the queries and then matched each one to one of 11,933 NMRA-approved medicines. Because users type imprecisely, matching was approximate: a fuzzy string-matching method based on Levenshtein distance, which measures how many character edits separate two strings, with a threshold of five or fewer edits.
Once queries were tied to medicines, the authors grouped them into therapeutic categories using a classification aligned with the Anatomical Therapeutic Chemical (ATC) system. They then looked at the resulting patterns — which categories and which individual medicines were searched most, how much of the catalogue was never searched at all, and which searches returned nothing. Separately, they measured system performance in terms of how often queries resolved and how long they took, and used processing times to estimate energy consumption.
The abstract does not specify how individual users or sessions were handled, what the platform's user base looks like, or how the "mismatch" against essential medicines lists was scored.
Why This Matters
Impact on research. The study treats search logs as a population-level data source for health demand, rather than as mere website telemetry, and argues that this kind of analysis is both feasible and underexplored in LMICs. The fuzzy-matching-plus-ATC pipeline it describes is a reusable template for anyone with messy real-world query data and a controlled drug vocabulary. It also adds an environmental dimension to digital health evaluation, estimating energy use alongside the more familiar latency and resolution metrics.
Real-world applications:
-
Pharmaceutical policy review. Search volumes for vitamins, antibiotics, anti-diabetes and antihypertensive medicines — and the divergence from essential medicines lists — could inform where information campaigns, regulatory attention, or formulary discussions are directed.
-
Health education and communication targeting. Knowing which therapeutic areas the public actually asks about can shape what public-facing medicines information is produced and prioritised.
-
Gap-filling from zero-result queries. Searches that resolve to nothing identify medicines or concepts the public wants but the platform cannot currently answer, pointing to content or catalogue gaps.
-
Digital health infrastructure planning. The latency, resolution, and energy figures offer a concrete benchmark for other LMIC health platforms weighing the cost of national-scale deployment.
Industry relevance. The reported performance profile — median 8 ms latency, over 99% resolution, roughly 0.12 kWh per million queries — is directly relevant to teams operating public-facing health information services and to anyone arguing that low-cost, low-energy architecture is achievable at national scale. The combination of approximate string matching over a fixed drug vocabulary with lightweight processing is a practical pattern for search systems in resource-constrained environments.
Future Directions
-
Extend the approach to other national platforms. Since the abstract frames LMIC medicine search behaviour as poorly characterised, the natural next step is replicating this analysis in other countries and testing whether the patterns — few dominant categories, heavy concentration, a large unqueried tail — generalise.
-
Turn the demand/policy mismatch into an actionable policy instrument. The abstract asserts that usage analytics could strengthen pharmaceutical policy but does not describe a mechanism, leaving open how divergence findings should be routed into regulatory or formulary processes, and how such interventions would be evaluated.
-
Investigate the zero-result and never-queried sets. The roughly 1% zero-result queries and the 40%–50% of registered medicines never queried are both reported but unexplained; follow-up work could determine which reflect genuine unmet need, which reflect poor discovery or naming, and which simply reflect low clinical importance.
-
Validate the environmental and performance estimates more deeply. Energy consumption was estimated from processing times rather than measured directly, so independent measurement of the full serving stack — and whether the reported efficiency holds under growing traffic — is an open question.
Target Audience
This paper is most useful to health informatics researchers and digital health platform operators, particularly those working in low- and middle-income countries; pharmaceutical policy analysts and medicines regulators interested in using real-world demand data; and engineers designing search infrastructure for national-scale health services where latency and energy cost matter. Public health and health economics researchers studying information-seeking behaviour will also find the log-based approach relevant. Readers looking for clinical outcomes, user-level demographics, or comparative evaluations against other platforms will not find them in the abstract.
Authors’ abstract
National-scale digital health platforms generate usage data that reveal population healthcare needs. MediVerify, Sri Lanka's online medicines information platform providing access to National Medicines Regulatory Authority (NMRA)-approved medicines, recorded over 1.49 million queries in its first year. Large-scale medicine search behaviour in low- and middle-income countries (LMICs) remains poorly characterised. Objectives: To characterise medicine information-seeking behaviour, identify mismatches between public demand and essential medicines policy, and evaluate technical performance. Methods: We retrospectively analysed 1,497,304 anonymised queries submitted between July 2024-2025. Queries were normalised and mapped to 11,933 approved medicines using fuzzy matching based on Levenshtein distance (threshold <=5). Therapeutic categories were assigned using an Anatomical Therapeutic Chemical (ATC)-aligned classification. Query patterns and system performance were analysed, and energy consumption was estimated from processing times. Results: Vitamins/minerals (11.74%), antibiotics (10.57%), anti-diabetes (7.02%), and antihypertensives (6.84%) dominated searches. The top 20 accounted for 37.3% of queries, while approximately 40%-50% of registered medicines were never queried. Zero-result queries (~ 1%) indicated unmet information needs. Frequently searched medicines diverged from essential medicines lists. Median latency was 8ms, with >99% query resolution. Estimated energy consumption was approximately 0.12 kWh per million queries. Conclusions: Large-scale medicine search data provide insights into healthcare information demand and policy alignment. MediVerify demonstrates that a national digital health platform can achieve high utilisation at low computational and environmental cost. Usage analytics could strengthen pharmaceutical policy and digital health infrastructure in LMICs.