Pharmaceutical · Pharmacovigilance and market intelligence
Drug safety signals found in text that no team could read in full
A global Swiss pharmaceutical company had to work through enormous volumes of multilingual medical texts, social media posts and patient feedback, looking for adverse drug events, medical entities and shifts in market sentiment. We built the NLP pipeline that automated what teams had been doing by hand, across five languages.
Key results
Client: Global Swiss pharmaceutical company (under NDA)
A multinational pharmaceutical enterprise operating across European and global markets. The company is under NDA, so it is not named here.
- Industry
- Pharmaceutical
- Use case
- Pharmacovigilance and text analysis
- AI approach
- NLP, named entity recognition, sentiment analysis
- Data source
- Medical notes, social media, online posts
- Languages
- DE, EN, FR, IT, ES
- Engagement
- Multi-phase delivery with a partner
In short
What was buried in the unstructured text
- Pharmacovigilance and market intelligence data sat inside millions of unstructured texts in five languages: patient feedback, notes from medical professionals, social media, regulatory documents.
- Teams reviewed and classified those texts by hand under strict accuracy and compliance requirements, and the work could not scale with the volume.
- The pipeline recognises medical entities, strips patient and physician data, detects adverse drug events in online posts and tracks sentiment around a launch.
- Every stage is GDPR-compliant by construction, because sensitive health data is what made automating this hard in the first place.
The starting point
The challenge
Critical pharmacovigilance and market intelligence data was locked inside millions of unstructured texts in five languages. Manual review could not keep up, and sensitive health information made automation non-trivial: every pipeline had to be GDPR-compliant from the start.
In pharmaceutical development and post-market surveillance the volume of text is enormous: patient feedback, notes from medical professionals, social media discussions, regulatory documents. At this company, teams were reviewing and classifying all of it by hand, across five languages, under strict accuracy and compliance requirements.
That was slow, it was inconsistent, and it did not scale.
Meanwhile the signals that matter, adverse drug events, a shift in public sentiment around a new drug, a mention of an unknown side effect, sat buried in unstructured text that no team could realistically process in full.
The build
What we built
Together with our partner we delivered a series of connected NLP capabilities, each closing a specific gap in the client’s text analysis workflow.
Recognising medical entities in five languages
Automated recognition of drug names, active substances, dosage forms and related terms across all five languages. Pharmaceutical nomenclature is its own problem: the system handles brand names, generics and the informal references people actually use in patient-generated content.
De-identification before anything else
Patient and physician data had to come out before any analysis could start. Automated anonymisation pipelines strip personally identifiable information from medical notes and records, which is what makes the downstream analysis both possible and GDPR-compliant.
Adverse events found in public posts
We developed an approach that identifies adverse events tied to specific drugs by analysing opinions and posts published online. That moves pharmacovigilance from reactive manual review to proactive monitoring, and it catches signals a manual process would miss entirely.
Sentiment around a launch
Sentiment tracking across social media channels shows how a drug is perceived before and during its market launch: how patients, caregivers and healthcare professionals are actually talking about it, rather than how a survey says they might.
Preprocessing built for real text
All of this had to survive five languages and the state social media text arrives in: abbreviations, colloquial language, medical slang, inconsistent formatting. We developed preprocessing and NLP algorithms that normalise and interpret those variations reliably.
What changed
The results
Before
Manual review of multilingual medical texts. Adverse events found reactively, if at all. No systematic social media monitoring, and GDPR compliance handled case by case.
After
Automated pipelines across five languages, proactive adverse event detection, real-time sentiment tracking, and de-identification built into every stage.
The company can now process medical data at a scale that manual workflows made impossible, and unknown adverse events surface from online sources automatically.
The frequency of known adverse events can be checked against real-world data instead of relying on clinical reports alone.
Market launch teams see how a drug is being discussed publicly as it happens, which shortens the time between an emerging concern and a response to it.
All of it runs inside a framework that respects patient privacy from the ground up. Automated de-identification means GDPR compliance is a property of every pipeline, not a step someone remembers to add.
Questions about this project
Pharmacovigilance and text analysis. Critical pharmacovigilance and market intelligence data was locked inside millions of unstructured texts in five languages. Manual review could not keep up, and sensitive health information made automation non-trivial: every pipeline had to be GDPR-compliant from the start.
German, English, French, Italian, Spanish: 5 langs. Adverse drug events found in unstructured online sources: Automated. Full de-identification of patient and physician data: GDPR. Market sentiment before and during a drug launch: Real time.
NLP, named entity recognition, sentiment analysis. Technology used: Natural Language Processing, Named Entity Recognition, Sentiment analysis, Text de-identification, Multilingual NLP, Adverse event detection, Social media analysis, Python.
Technology used
Talk to the people who built this
theBlue.ai comes back within one business day.
More case studies

Market research · InSaaS.ai
Anonymisation built into the product rather than bolted on after
Read more
Healthcare · Tirol Kliniken Innsbruck
Patient data anonymised on the hospital’s own servers
Read more
Telecommunication · Arvato Bertelsmann