AI Architecture & Integration · Local LLMs
Local LLMs in Practice: How We Use Them at theBlue.ai
For many AI tasks it is cheaper, safer and easier to control an open-source model on your own infrastructure than to send data to a cloud API. Three systems from our daily work show where local models play to their strengths and where they reach their limits.
Key takeaways
- Local LLMs earn their place on volume, data sensitivity or both. Without those requirements, a cloud API is usually the honest choice.
- The task decides the model: Gemma 4 26B where users wait (88 tokens per second), the larger 31B variant for background jobs where accuracy counts.
- Quantization, an MTP layer, mixture-of-experts and the serving stack can more than double throughput on the same GPUs.
- Running locally removes the risk of external data egress. In exchange, application security, access control and operations sit with your own team.
A default approach has emerged around large language models (LLMs): whenever a task calls for AI, the data goes to ChatGPT or a similar cloud tool. Often that is the right choice. For a whole class of tasks, though, running LLMs locally, meaning an open-source model on infrastructure you control, turns out to be cheaper, safer and easier to control.
This view is based on systems we have built ourselves and use every day. Below we describe what that looks like in practice, where local models offer a real advantage and where they reach their limits.
What is a local LLM?
A local LLM is an open-source model, for example from the Gemma, Llama or Qwen families, that runs on your own hardware rather than on an external provider’s server. For a business, two consequences matter. First, data never leaves the company’s infrastructure. Second, there is no per-query fee. Together, these two factors change both the cost profile and the risk profile of many applications.
Three systems, three lessons
We run several such systems. Here we describe three of them, starting with the one that best shows what it means to fit a model to a task.
1.Tender portal: screening public procurement at scale
An internal portal automatically searches for public tenders in our areas of interest and downloads everything available about them: descriptions, requirements, attachments and supporting documentation. After pre-filtering, the local model Gemma 4 26B (26 billion parameters) with an MTP (multi-token prediction) layer takes over. It handles two stages:
- Eligibility check: the model reads the header of each tender, meaning title, buyer, category and attachment names, and gives a short verdict on whether it fits our company profile, with a confidence level and a brief explanation. Only relevant tenders remain on the list afterwards.
- Documentation analysis: once documentation and attachments are downloaded, the model returns a structured summary: contract, financial, staffing and experience requirements, required certificates, the amount and form of the bid bond, key dates, risks such as contractual penalties, and a recommendation (RECOMMEND, CONSIDER or REJECT) with justification.
The scale is real. In a single day, the model processes 15 to 20 million tokens, roughly 15,000 A4 pages. Over a year, that adds up to 5.5 to 7.3 billion tokens. Running the same workload through GPT-5 would cost $6,875 to $9,125 for input tokens alone.
Calculated from OpenAI’s official pricing as of July 2026
The model choice was deliberate. In the portal, a user is actively waiting for the result, so response time matters. On our hardware, Gemma 4 26B generates around 88 tokens per second, the larger 31B variant 40. The 31B version is slightly more accurate, but with user experience in mind the 26B model is the right compromise here. It owes its speed to the MTP layer and the mixture-of-experts (MoE) architecture, both described below. It also has a large context window, which lets us analyse an entire tender with its attachments in one pass, and it handles Polish well.
Weighing accuracy, latency, context window, language support and hardware footprint against each other leads to a clear result in this case: large volumes of tenders can be processed without worrying about API costs or query limits.
2.Policy Insider: where per-query pricing breaks down
Policy Insider is an AI-powered policy and regulatory intelligence platform. It helps public affairs, government relations and strategic advisory teams follow policy and regulatory developments at scale, without manual tracking, spreadsheets or fragmented tools. Here the local LLM Gemma 4 31B analyses content from a large number of policy-related websites, about 28 million tokens per day. For each site, the model prepares the content, summarises it and flags whether it matches a topic. Relevant sites feed into reports, newsletters and heatmaps that show where change is most intense.
Here we deliberately use the larger, more accurate model with 31 billion parameters. Policy Insider processes data in cyclical background jobs, so it makes no difference whether a batch takes one minute or thirty. What counts is classification accuracy, because the quality of the reports depends on it. The principle stays the same: the task decides the model.
This is also a case where a paid API stops making economic sense. Millions of tokens per cycle mean high, constantly growing costs and ongoing trouble with query limits. With a local LLM, each additional website costs practically nothing: the hardware is paid for once, and volume is no longer a financial constraint. Our Policy Insider case study shows more of the platform.
3.Internal chatbot with RAG: data that stays inside the company
Cost and scale are one side, data control the other. The third system serves our own team: a chatbot that searches our internal documentation and answers questions. Instead of digging through dozens of documents, you ask a question in natural language and get an answer with a pointer to the source. The design follows the retrieval-augmented generation (RAG) pattern: the system first finds matching passages in our documents, and the model composes the answer from them, drawing on those passages over knowledge from its training.
The system runs locally on our own GPUs, this time with a model from outside the Gemma family: Qwen3 30B. This shows the principle behind all three implementations most clearly: once the task decides the model, there is no reason to stay within one model family. Qwen3 handles conversation well, copes with long documents and stays fast despite its nominal 30 billion parameters, because each query activates only a fraction of them (mixture-of-experts). It therefore stays responsive when many people ask at once and can draw on large parts of the documentation for every answer.
What matters most is where the work happens: our internal knowledge, meaning projects, proposals and know-how, never goes to an external provider. For us that means peace of mind. For clients in regulated sectors such as finance, insurance, government or healthcare, it can be the precondition for even discussing AI adoption.
The three systems at a glance
| System | Model | What matters |
|---|---|---|
| Tender portal | Gemma 4 26B with MTP | Response time, 15 to 20M tokens a day |
| Policy Insider | Gemma 4 31B | Accuracy, about 28M tokens a day |
| Internal chatbot | Qwen3 30B (MoE) | Data stays internal, many users at once |
Limits and trade-offs
Local LLMs do not fit every case. These limits are worth knowing:
- Upfront cost: hardware is a capital investment. It pays off at sufficient volume and over a long enough period. For small scale or occasional use, a cloud API can be cheaper.
- Hardware as a ceiling: the model runs on your own machines and is limited by their compute, memory and architecture. The largest models may simply not run on a given setup. A cloud API, by contrast, can be called from any old laptop with an internet connection, because the compute sits with the provider.
- Quality gap on the hardest tasks: for complex problems that require deep reasoning, leading commercial models still often deliver better results. Not every task needs the highest tier of reasoning, though, so choosing the model to fit pays off.
- Ongoing operations: the model has to be hosted, monitored and updated. That is a permanent operational responsibility.
How to run LLMs locally: what to look for
How well a local deployment works depends as much on how the model is served as on which model you pick. A few technical decisions have a large impact on cost, quality and speed.
- Quantization: the model’s weights can be stored at lower precision, for example 8-bit or 4-bit instead of the default. That saves GPU memory and usually speeds up inference. The more aggressive the quantization, the bigger the savings and the larger the possible drop in quality. The goal is the level at which the model runs faster and uses less memory while quality stays good enough for the task. That is why we validate every level with quality tests on our own data.
- MTP (multi-token prediction) layer: a standard model generates its answer token by token, which caps its speed. An MTP layer predicts several tokens in one step and speeds up generation without hurting quality. In our test it raised the portal model’s throughput from around 72 to 88 tokens per second, and the 31B variant’s from 16 to 40, more than double. However, MTP is not available for every model, and the layer makes the model slightly larger.
- Mixture-of-experts (MoE): a dense model uses every parameter for every token. An MoE model is split into many specialised sub-networks, the experts, and a small router picks only a few of them per token. On paper the model has tens of billions of parameters, but a query activates only a fraction. An MoE model therefore runs much faster and cheaper than a dense model of the same size while keeping most of the quality. The model in the tender portal activates only about 4 of its 26 billion parameters per query. The catch is on the hardware side: all experts still have to fit into GPU memory.
- Context window and cache: the longer the text passed to the model, for example an entire tender with attachments, the more GPU memory the cache takes up, where the model stores the context it has already processed (the KV cache). How much context is usable in practice also depends on how much memory remains after loading the weights. This has to be planned from the start.
- Serving software: the serving stack determines how much throughput the same hardware delivers. vLLM and SGLang are built for high throughput and many parallel queries and suit production use. Ollama and llama.cpp are quicker to set up and work well for prototypes or smaller deployments, but lose performance under load.
- Many parallel queries: when many people or processes use the model at the same time, what matters is grouping parallel queries into efficient batches (continuous batching). Mature serving software does this automatically, which is why the same hardware serves several times the queries of purely sequential processing.
- Hardware selection: the bottleneck is usually GPU memory. It has to hold the weights, the context window and headroom for parallel queries. If the hardware can be chosen freely, planning starts with calculating how much memory the desired model needs at the intended quantization level and expected load.
These decisions are connected. Quantization, serving stack and hardware interact, and a mistake in one place usually shows up as poor performance in another.
When a local model makes sense, and when it does not
Each of the three systems hits the point where running the model locally is the better option. Just as many cases call for a cloud API.
A local model usually makes sense when
- data cannot leave the company: personal data under the GDPR, medical records, financial information, trade secrets or documents subject to sector rules (finance, insurance, healthcare, public administration) often cannot go to an external provider at all, or only with safeguards that are more effort than running the model in-house. In regulated sectors, local deployment is frequently the precondition for the project.
- volume is high and predictable: at millions of tokens per day in a recurring pattern, API costs quickly exceed the amortised cost of your own hardware. The tender portal and Policy Insider both fall here.
- you need full control over the model: fine-tuning on your own data, a frozen model version whose behaviour does not change after provider updates, custom quantizations or non-standard serving setups are impossible or heavily restricted with a closed API.
A cloud API usually makes sense when
- volume is low or unpredictable: with occasional use, hardware amortisation never catches up with per-query pricing.
- you need the highest reasoning quality: on the hardest tasks, commercial frontier models are still ahead. If quality on the toughest queries decides the value of the whole system, a cloud model earns its price there. Local models are catching up, though, and Kimi K3 is a good example.
- you lack staff for operations: someone has to keep the GPUs running, update the serving stack and respond to incidents. Without that capacity in-house, an API is the most honest choice.
A note on data security
The security argument for local models reaches further than “your data does not leave the building” and is more nuanced at the same time.
A local deployment keeps data inside infrastructure you control, with your access policies, your logging and your retention rules. There are no third-party terms of service to reconcile with your own compliance obligations, no uncertainty about whether queries will later be used for training, and no dependency on the provider’s security posture. In regulated sectors this often shortens the compliance discussion from months to days.
A model on your own hardware is still not secure by default. It is only as secure as its environment. Access has to be authenticated and logged, prompts and outputs can contain sensitive data, and RAG systems inherit the permissions of the document store: a chatbot that indexes files a user should not see will happily reveal their contents.
A local model removes a large category of risk, external data egress, and exchanges it for a smaller, more familiar one: application and infrastructure security. For most organisations that is a good trade, provided they take the second category seriously.
Conclusion: local LLMs are a question of the task
Local LLMs are a long-term operational commitment. Where volume is high or data must not leave the company, they are often the more economical and sometimes the only permissible solution. How the model is served matters as much as which model you choose.
At theBlue.ai we deploy local LLMs for clients and inside our own company. From these projects we know where the real trade-offs lie and when a local model, a cloud API or a hybrid setup is the right answer. Find out more on our LLM development page.
Local or cloud for your use case?
In the AI discovery workshop we answer exactly that question within one to three weeks and give you an architecture proposal you can act on.
Request the workshopFrequently asked questions
A local LLM is an open-source model, for example from the Gemma, Llama or Qwen families, that runs on your own hardware rather than on an external provider’s server. Data does not leave the company, and there is no per-query fee.
When data must not leave the company, when volume is high and predictable, for example millions of tokens per day, or when you need full control over the model. With low or fluctuating volume, the highest reasoning requirements or no staff for operations, a cloud API is usually better.
Our tender portal processes 15 to 20 million tokens a day, 5.5 to 7.3 billion a year. Through GPT-5 that would cost $6,875 to $9,125 for input tokens alone, based on OpenAI pricing from July 2026. Locally, the hardware is paid for once.
No. It removes the risk of external data egress, but it is only as secure as its environment. Access has to be authenticated and logged, and RAG systems inherit the permissions of the document store.
Aleksandra Osztynowicz