AI Architecture & Integration · LLMOps

LLM Observability and Monitoring: Keeping AI Applications in View

An LLM application that impresses in testing can quietly get worse in production: answers drift, costs rise, users drop off. Observability shows what happens in every request and gives you the basis to manage quality and cost.

Aleksandra Osztynowicz 5 min read
Illustration of a man in a suit holding two magnifying glasses, with question and exclamation marks above him

Key takeaways

  • LLM observability captures input, context, model response, latency, cost and rating for every request.
  • Monitoring tells you that something is wrong. Observability helps you understand why.
  • The main measures are answer quality, cost per request, latency, errors and user feedback.
  • Tools such as Langfuse and LangSmith combine tracing, evaluation and prompt management. Langfuse can also be self-hosted.

What LLM observability is

LLM applications behave differently from classic software. The same question can produce different answers, a new prompt or model changes the results, and an error often shows up as a wrong statement instead of an error message. LLM observability makes these processes traceable. It relies on three kinds of data:

  • Logs: chronological records of events, such as inputs, answers and errors.
  • Metrics: measurements such as latency, token usage, cost and error rates.
  • Traces: the full path of a request through the system, from document retrieval and the model call to tools and agent steps.

Monitoring and observability

Monitoring tracks defined metrics and raises an alert when thresholds are crossed, for example when latency or cost climbs. Observability goes one level deeper: in the individual trace it shows which step caused the poor answer, such as a wrongly retrieved document or an outdated prompt. Production needs both.

What is measured in production

AreaTypical questions
Answer qualityDo answers match the sources? How often is something invented or left out?
CostHow many tokens does a request use, and which steps drive the cost?
LatencyHow long do users wait, and which step slows things down?
ErrorsWhere do calls fail, and where do tools or interfaces return nothing?
User feedbackWhich answers are flagged as wrong, and why?
VersionsWhich prompt and which model produced an answer?

Spot checks are rarely enough for quality. Fixed test sets with real questions work well, re-evaluated automatically whenever the prompt, model or knowledge base changes.

In practice

In practice: apoQlar

The RAG assistant for security questionnaires uses LangFuse to track user feedback together with prompt versions, cost and latency. When users flag an answer as inaccurate, it often points to a gap in the documentation, so the teams know which policy to update.

Read the apoQlar case study

In practice: Policy-Insider.AI

The platform summarises policy documents from several EU languages automatically. To keep quality steady in production, we built an iterative evaluation pipeline that continuously monitors inconsistent output and hallucination risk.

Read the Policy-Insider.AI case study

In practice: technical inspection authority

The prototype of the citizen chatbot with eleven domain agents runs on Microsoft Azure with Entra ID authorisation and Langfuse monitoring. That shows at any time how the system is performing, in line with the authority’s security requirements.

Read the citizen chatbot case study

Langfuse and LangSmith

Specialised tools take on much of this work. Two widely used examples:

  • Langfuse is open source and offers tracing, evaluations and prompt management. It can be used as a cloud service or self-hosted in your own infrastructure with Docker or Kubernetes. Some add-on features require a licence there.
  • LangSmith comes from LangChain and covers tracing, evaluation and prompt engineering. It also works with other frameworks and model providers and runs in the cloud, hybrid or self-hosted.

When choosing, what matters most is where log data may be stored, whether personal data can be masked before storage, how well the tool connects to existing systems and frameworks, and whether evaluations can run automatically.

How to introduce observability

  1. Log from the start: build tracing into the prototype so you have baseline figures from day one.
  2. Define quality: decide what makes a good answer and build a test set of real questions.
  3. Make cost and latency visible: per request and per step, so expensive spots show up early.
  4. Collect feedback: users can rate answers easily, and the ratings feed into evaluation.
  5. Safeguard changes: every new prompt and every new model is checked against the test set before use.

How operating AI models is organised as a whole is covered in our article Ensuring reliable and scalable AI solutions with MLOps.

Securing your LLM application in production?

We set up tracing, evaluation and cost control with you, in the cloud or in your own infrastructure.

Request a process analysis

Frequently asked questions

The ability to understand how an LLM application behaves in production. Logs, metrics and traces are captured for every request: input, context, answer, latency, cost and rating.

Monitoring tracks defined metrics and reports deviations. Observability shows in the individual flow why an answer was poor or a request was expensive.

Mainly answer quality, cost per request, latency, errors, user feedback and the versions of prompt and model.

Yes. Langfuse is open source and can be self-hosted with Docker or Kubernetes. Some add-on features require a licence.

Aleksandra Osztynowicz

About the author

Aleksandra Osztynowicz

AI Engineer, theBlue.ai

Aleksandra has been building custom AI solutions at theBlue.ai since 2021, with a focus on agentic implementations and local LLM deployments that precisely fit enterprise needs. As an AI Engineer, she helps organizations automate their processes with production-grade systems, from on-premise open-source models to RAG-based internal knowledge bases for regulated industries, and continuously expands her knowledge in the rapidly changing world of artificial intelligence.

In her articles, she shares practical experience from real enterprise AI projects and shows that deploying AI in companies doesn’t need to be complicated.