Case studies

Pharmaceutical · Pharmacovigilance and market intelligence

Drug safety signals found in text that no team could read in full

A global Swiss pharmaceutical company had to work through enormous volumes of multilingual medical texts, social media posts and patient feedback, looking for adverse drug events, medical entities and shifts in market sentiment. We built the NLP pipeline that automated what teams had been doing by hand, across five languages.

Key results

5 langs
German, English, French, Italian, Spanish
Automated
Adverse drug events found in unstructured online sources
GDPR
Full de-identification of patient and physician data
Real time
Market sentiment before and during a drug launch

Client: Global Swiss pharmaceutical company (under NDA)
A multinational pharmaceutical enterprise operating across European and global markets. The company is under NDA, so it is not named here.

Industry
Pharmaceutical
Use case
Pharmacovigilance and text analysis
AI approach
NLP, named entity recognition, sentiment analysis
Data source
Medical notes, social media, online posts
Languages
DE, EN, FR, IT, ES
Engagement
Multi-phase delivery with a partner

In short

What was buried in the unstructured text

  • Pharmacovigilance and market intelligence data sat inside millions of unstructured texts in five languages: patient feedback, notes from medical professionals, social media, regulatory documents.
  • Teams reviewed and classified those texts by hand under strict accuracy and compliance requirements, and the work could not scale with the volume.
  • The pipeline recognises medical entities, strips patient and physician data, detects adverse drug events in online posts and tracks sentiment around a launch.
  • Every stage is GDPR-compliant by construction, because sensitive health data is what made automating this hard in the first place.

The starting point

The challenge

Critical pharmacovigilance and market intelligence data was locked inside millions of unstructured texts in five languages. Manual review could not keep up, and sensitive health information made automation non-trivial: every pipeline had to be GDPR-compliant from the start.

In pharmaceutical development and post-market surveillance the volume of text is enormous: patient feedback, notes from medical professionals, social media discussions, regulatory documents. At this company, teams were reviewing and classifying all of it by hand, across five languages, under strict accuracy and compliance requirements.

That was slow, it was inconsistent, and it did not scale.

Meanwhile the signals that matter, adverse drug events, a shift in public sentiment around a new drug, a mention of an unknown side effect, sat buried in unstructured text that no team could realistically process in full.

The build

What we built

Together with our partner we delivered a series of connected NLP capabilities, each closing a specific gap in the client’s text analysis workflow.

01

Recognising medical entities in five languages

Automated recognition of drug names, active substances, dosage forms and related terms across all five languages. Pharmaceutical nomenclature is its own problem: the system handles brand names, generics and the informal references people actually use in patient-generated content.

02

De-identification before anything else

Patient and physician data had to come out before any analysis could start. Automated anonymisation pipelines strip personally identifiable information from medical notes and records, which is what makes the downstream analysis both possible and GDPR-compliant.

03

Adverse events found in public posts

We developed an approach that identifies adverse events tied to specific drugs by analysing opinions and posts published online. That moves pharmacovigilance from reactive manual review to proactive monitoring, and it catches signals a manual process would miss entirely.

04

Sentiment around a launch

Sentiment tracking across social media channels shows how a drug is perceived before and during its market launch: how patients, caregivers and healthcare professionals are actually talking about it, rather than how a survey says they might.

05

Preprocessing built for real text

All of this had to survive five languages and the state social media text arrives in: abbreviations, colloquial language, medical slang, inconsistent formatting. We developed preprocessing and NLP algorithms that normalise and interpret those variations reliably.

Drug safety signals found in text that no team could read in full

What changed

The results

Before

Manual review of multilingual medical texts. Adverse events found reactively, if at all. No systematic social media monitoring, and GDPR compliance handled case by case.

After

Automated pipelines across five languages, proactive adverse event detection, real-time sentiment tracking, and de-identification built into every stage.

The company can now process medical data at a scale that manual workflows made impossible, and unknown adverse events surface from online sources automatically.

The frequency of known adverse events can be checked against real-world data instead of relying on clinical reports alone.

Market launch teams see how a drug is being discussed publicly as it happens, which shortens the time between an emerging concern and a response to it.

All of it runs inside a framework that respects patient privacy from the ground up. Automated de-identification means GDPR compliance is a property of every pipeline, not a step someone remembers to add.

Questions about this project

Pharmacovigilance and text analysis. Critical pharmacovigilance and market intelligence data was locked inside millions of unstructured texts in five languages. Manual review could not keep up, and sensitive health information made automation non-trivial: every pipeline had to be GDPR-compliant from the start.

German, English, French, Italian, Spanish: 5 langs. Adverse drug events found in unstructured online sources: Automated. Full de-identification of patient and physician data: GDPR. Market sentiment before and during a drug launch: Real time.

NLP, named entity recognition, sentiment analysis. Technology used: Natural Language Processing, Named Entity Recognition, Sentiment analysis, Text de-identification, Multilingual NLP, Adverse event detection, Social media analysis, Python.

Technology used

Natural Language Processing Named Entity Recognition Sentiment analysis Text de-identification Multilingual NLP Adverse event detection Social media analysis Python

Talk to the people who built this

theBlue.ai comes back within one business day.

Contact us