v3.2 April 2026

Lumen

Inference, evaluation, and observability, quietly engineered.

The single audited surface for running, measuring, and shipping language model workloads. Built for engineering teams who treat AI as infrastructure, not as a marketing department.

  • CLAUDE-3.5-SONNET

    412ms

  • GPT-4O-MINI

    218ms

  • LLAMA-3.1-70B

    384ms

  • MISTRAL-LARGE

    502ms

  • GEMINI-1.5-PRO

    326ms

  • DEEPSEEK-V3

    287ms

As shipped — v3.2.1

The console — everything you ship, on one surface

dashboard with inference overview.

§ 01

The Index

Four primitives. Nothing more.

We don't believe in forty-feature dashboards or AI that does your laundry. Lumen does four things — and does them with the rigor your production workloads quietly deserve.

N° 01

Inference, routed

— OpenAI · Anthropic · self-hosted

N° 01

Inference, routed

— OpenAI · Anthropic · self-hosted

N° 02

Evaluations, versioned

— LLM-as-judge · regression suites

N° 02

Evaluations, versioned

— LLM-as-judge · regression suites

N° 03

Observability, auditable

— full traces · cost · latency

N° 03

Observability, auditable

— full traces · cost · latency

N° 04

Datasets, annotated

— from production traces

N° 04

Datasets, annotated

— from production traces

§ 02

Specimen

A specimen of the type.

Each primitive composed and laid out as in a type specimen — numerals, ascenders, ligatures. The mark of well-set software.

01

Inference, routed.

One SDK across twenty-three providers. Caching, retries, and fallbacks arrive as defaults — not as features wired in over a long weekend.

Provider routing

23 supported

Streaming

SSE · WS

Fallbacks

automatic

Self-hosted models

vLLM · TGI

02

Evaluations, versioned.

Every regression caught before customers find it. Suites build themselves from real traces, prompts and models compared side by side.

LLM-as-judge

built-in

Custom assertions

unlimited

CI integration

GitHub · GitLab

Versioned suites

git-backed

03

Observability, auditable.

Nothing about your system is opaque, including its costs. Filter, replay, export every trace; tokens, latency attach themselves to each request.

Trace retention

90 days

SOC 2 Type II

audited

Self-hosted option

included

Data residency

EU · US · APAC

04

Datasets, annotated.

Every production trace becomes labeled data. Curate, version, and export evaluation sets without leaving the surface that captured them.

Source

production traces

Labeling

inline · API

Export formats

JSONL · CSV

Versioning

git-backed

§ 03

Testimonials

"We replaced four internal tools with Lumen and now ship faster and effective with a smaller team."

Elena VukšićStaff Engineer · Halcyon Labs

"We replaced four internal tools with Lumen and now ship faster and effective with a smaller team."

Elena VukšićStaff Engineer · Halcyon Labs

Built for teams who take their infrastructure seriously.

Start with the free tier. Scale when you're ready. No credit card, no calls with sales, no nonsense. The way it should be.

Lumen

"Software that earns its place not through novelty, but through quiet competence."

Lumen is the audited surface for running, measuring, and shipping language model workloads — built for engineering teams who treat AI as infrastructure.

All systems operational

99.98%

— Last 30 days

Measured continuously across regions. Independently audited. No exceptions, no asterisks.

— Regions

EU · US · APAC

Lumen

.

© MMXXVI · LUMEN LABS, D.O.O.

Lumen

"Software that earns its place not through novelty, but through quiet competence."

Lumen is the audited surface for running, measuring, and shipping language model workloads — built for engineering teams who treat AI as infrastructure.

All systems operational

99.98%

— Last 30 days

Measured continuously across regions. Independently audited. No exceptions, no asterisks.

— Regions

EU · US · APAC

Lumen

.

© MMXXVI · LUMEN LABS, D.O.O.

Lumen

"Software that earns its place not through novelty, but through quiet competence."

Lumen is the audited surface for running, measuring, and shipping language model workloads — built for engineering teams who treat AI as infrastructure.

All systems operational

99.98%

— Last 30 days

Measured continuously across regions. Independently audited. No exceptions, no asterisks.

— Regions

EU · US · APAC

Lumen

.

© MMXXVI · LUMEN LABS, D.O.O.

Create a free website with Framer, the website builder loved by startups, designers and agencies.