AI Architecture & Integration · LLMOps
LLM Observability and Monitoring: Keeping AI Applications in View
An LLM application that impresses in testing can quietly get worse in production: answers drift, costs rise, users drop off. Observability shows what happens in every request and gives you the basis to manage quality and cost.
Key takeaways
- LLM observability captures input, context, model response, latency, cost and rating for every request.
- Monitoring tells you that something is wrong. Observability helps you understand why.
- The main measures are answer quality, cost per request, latency, errors and user feedback.
- Tools such as Langfuse and LangSmith combine tracing, evaluation and prompt management. Langfuse can also be self-hosted.
What LLM observability is
LLM applications behave differently from classic software. The same question can produce different answers, a new prompt or model changes the results, and an error often shows up as a wrong statement instead of an error message. LLM observability makes these processes traceable. It relies on three kinds of data:
- Logs: chronological records of events, such as inputs, answers and errors.
- Metrics: measurements such as latency, token usage, cost and error rates.
- Traces: the full path of a request through the system, from document retrieval and the model call to tools and agent steps.
Monitoring and observability
Monitoring tracks defined metrics and raises an alert when thresholds are crossed, for example when latency or cost climbs. Observability goes one level deeper: in the individual trace it shows which step caused the poor answer, such as a wrongly retrieved document or an outdated prompt. Production needs both.
What is measured in production
| Area | Typical questions |
|---|---|
| Answer quality | Do answers match the sources? How often is something invented or left out? |
| Cost | How many tokens does a request use, and which steps drive the cost? |
| Latency | How long do users wait, and which step slows things down? |
| Errors | Where do calls fail, and where do tools or interfaces return nothing? |
| User feedback | Which answers are flagged as wrong, and why? |
| Versions | Which prompt and which model produced an answer? |
Spot checks are rarely enough for quality. Fixed test sets with real questions work well, re-evaluated automatically whenever the prompt, model or knowledge base changes.
In practice
In practice: apoQlar
The RAG assistant for security questionnaires uses LangFuse to track user feedback together with prompt versions, cost and latency. When users flag an answer as inaccurate, it often points to a gap in the documentation, so the teams know which policy to update.
In practice: Policy-Insider.AI
The platform summarises policy documents from several EU languages automatically. To keep quality steady in production, we built an iterative evaluation pipeline that continuously monitors inconsistent output and hallucination risk.
In practice: technical inspection authority
The prototype of the citizen chatbot with eleven domain agents runs on Microsoft Azure with Entra ID authorisation and Langfuse monitoring. That shows at any time how the system is performing, in line with the authority’s security requirements.
Langfuse and LangSmith
Specialised tools take on much of this work. Two widely used examples:
- Langfuse is open source and offers tracing, evaluations and prompt management. It can be used as a cloud service or self-hosted in your own infrastructure with Docker or Kubernetes. Some add-on features require a licence there.
- LangSmith comes from LangChain and covers tracing, evaluation and prompt engineering. It also works with other frameworks and model providers and runs in the cloud, hybrid or self-hosted.
When choosing, what matters most is where log data may be stored, whether personal data can be masked before storage, how well the tool connects to existing systems and frameworks, and whether evaluations can run automatically.
How to introduce observability
- Log from the start: build tracing into the prototype so you have baseline figures from day one.
- Define quality: decide what makes a good answer and build a test set of real questions.
- Make cost and latency visible: per request and per step, so expensive spots show up early.
- Collect feedback: users can rate answers easily, and the ratings feed into evaluation.
- Safeguard changes: every new prompt and every new model is checked against the test set before use.
How operating AI models is organised as a whole is covered in our article Ensuring reliable and scalable AI solutions with MLOps.
Securing your LLM application in production?
We set up tracing, evaluation and cost control with you, in the cloud or in your own infrastructure.
Request a process analysisFrequently asked questions
The ability to understand how an LLM application behaves in production. Logs, metrics and traces are captured for every request: input, context, answer, latency, cost and rating.
Monitoring tracks defined metrics and reports deviations. Observability shows in the individual flow why an answer was poor or a request was expensive.
Mainly answer quality, cost per request, latency, errors, user feedback and the versions of prompt and model.
Yes. Langfuse is open source and can be self-hosted with Docker or Kubernetes. Some add-on features require a licence.
Aleksandra Osztynowicz