Case studies

Professional services · Data privacy and GDPR

Personal data stripped out before the analysis ever sees it

InSaaS.ai processes large volumes of text from social media, online forums and customer sources for market research and product planning. Every dataset can carry personally identifiable information that has to go before analysis, and reviewing it by hand was impossible at that scale. We integrated ShareMedix as an anonymisation API that detects and masks PII across their entire pipeline.

Key results

GDPR
Full compliance with European data privacy regulation
API
Integrated directly into the existing data pipeline
Parallel
Many concurrent anonymisation requests at scale
Auto
Names, addresses, phone numbers and IBANs masked

Client: InSaaS.ai
A data analytics company specialising in insights for marketing, market research and product development.

Industry
Professional services and market research
Use case
PII anonymisation in text data
AI approach
NLP-based PII detection
Data source
Social media, forums, customer data
Integration
API with parallel processing
Product
ShareMedix (theBlue.ai)

In short

Why the data could not be used as it arrived

  • InSaaS.ai’s work depends on analysing social media posts, forum discussions, customer feedback and internal datasets, and all of it carries personal data.
  • Under GDPR that data cannot be analysed in its raw form. Every piece of PII has to be found and removed first, and the volume and variety made a manual process impossible.
  • ShareMedix now sits in the pipeline as an API: data flows in, anonymised data flows out, with no manual step and no separate interface.
  • It also unlocked sources that had been too risky to use at all, which widened the range of insights InSaaS.ai can deliver.

The starting point

The challenge

InSaaS.ai had to make vast amounts of text GDPR-compliant before any of it could be analysed, and the volume and variety of personally identifiable information made manual anonymisation impossible at their scale.

InSaaS.ai builds analytics for marketing, market research and product development, and that work depends on large volumes of text: social media posts, forum discussions, customer feedback and internal datasets.

The insight in that data is valuable. The data itself is full of names, addresses, phone numbers, IBAN numbers and other sensitive details.

Under GDPR none of it can be processed in raw form, so every piece of PII has to be detected and removed before it enters the analytics pipeline. Manually that was not an option: too much data, too many kinds of PII, and regulatory stakes too high for a human process.

The build

What we built

We integrated ShareMedix, our NLP-powered anonymisation engine, directly into InSaaS.ai’s processing pipeline through an API. PII is detected and masked before the data reaches the analytics layer.

01

Finding personal data in text that was never structured

ShareMedix identifies names, addresses, phone numbers, IBANs, email addresses and other identifiers using NLP. It works on data from very different sources: a social media post written in informal language, a structured customer record, and everything between.

02

An API, not a tool someone has to open

Rather than a standalone application requiring manual operation, ShareMedix is an API InSaaS.ai calls from inside their existing pipeline. Data goes in, anonymised data comes out, and nobody has to remember to run it.

03

It does not become the bottleneck

Market research datasets can be enormous. The API handles many concurrent requests, so large volumes move through without anonymisation slowing the data preparation down.

04

Rules that can be adjusted

Different sources and use cases need different handling. ShareMedix supports white lists, black lists and configurable rules, so InSaaS.ai can tune what gets masked and what stays, and adjust as regulation and data sources change.

Personal data stripped out before the analysis ever sees it

What changed

The results

Before

PII scattered through massive text datasets. Manual review impossible at that scale, GDPR compliance hard to guarantee, and data left unused because of the privacy risk.

After

Automated detection and masking through the API. Compliance built into the pipeline, full datasets available for analysis, and no manual step anywhere.

InSaaS.ai can process and analyse their entire data volume with compliance maintained automatically. The anonymisation step went from a manual blocker to an invisible part of the pipeline.

It also unlocked data that had been too risky to use. Customer data and other sensitive sources that would have needed extensive manual review now go through automatically.

That widens the scope of what InSaaS.ai can tell their clients, which is the part a compliance measure is not usually expected to do.

Questions about this project

PII anonymisation in text data. InSaaS.ai had to make vast amounts of text GDPR-compliant before any of it could be analysed, and the volume and variety of personally identifiable information made manual anonymisation impossible at their scale.

Full compliance with European data privacy regulation: GDPR. Integrated directly into the existing data pipeline: API. Many concurrent anonymisation requests at scale: Parallel. Names, addresses, phone numbers and IBANs masked: Auto.

NLP-based PII detection. Technology used: Natural Language Processing, Named Entity Recognition, PII detection, API integration, ShareMedix platform, GDPR compliance.

Technology used

Natural Language Processing Named Entity Recognition PII detection API integration ShareMedix platform GDPR compliance

Talk to the people who built this

theBlue.ai comes back within one business day.

Contact us