Need help with your APIs? I offer API discovery, governance & evangelism services. Explore services →
API Evangelist API Evangelist
Discovery
Learnings
Guidance
Toolbox
Alignment
API Evangelist LLC
LangWatch website screenshot

LangWatch

LangWatch is an open-source LLM observability, evaluation, and AI agent testing platform. Built around OpenTelemetry-native tracing, LangWatch lets teams instrument LLM applications (LangChain, LangGraph, DSPy, OpenAI Agents, LiteLLM, Pydantic AI, CrewAI, AWS Bedrock, and more), run real-time and batch evaluations, version and deploy prompts, simulate multi-turn agent conversations against scripted scenarios and Judge Agents, manage labeled datasets, and govern model traffic through a virtual-key AI Gateway with budgets and semantic caching. The platform exposes a REST API at app.langwatch.ai, ships Python and TypeScript SDKs plus an MCP server, publishes the companion `scenario` agent-testing framework and `better-agents` standards, and runs on Apache-2.0 core (with `ee/` enterprise modules under commercial license) — deployable as LangWatch Cloud or self-hosted via Docker Compose, Helm, Kind, or full Kubernetes.

agent ready

Real signal across most facets with visible, nameable gaps — the contract exists but is thin, or the portal is good while governance and commercial terms are absent.

Kin Score

API Evangelist profiles LangWatch the way a machine reads it — 54 machine-readable artifacts across 28 APIs, pulled from the provider's own public surface and indexed so a developer, an analyst, or an AI agent can evaluate it against every other provider on the network.

Every provider in the network is reduced to the same set of machine-readable artifacts — OpenAPI contracts, event specifications, GraphQL schemas, runnable collections, pricing and rate-limit signals, security posture, OAuth scopes, and the agent surfaces (MCP servers and skills) that let software drive the API on its own. We profile them because the interface is the part of a company you can actually inspect: it is a truer signal of what a provider does than any marketing page. From those artifacts we compute the Kin Score — LangWatch scores 54.6/100 (developing), with a separate agent-readiness read of 48/100 (agent ready). The full breakdown is below, followed by every artifact we hold — each card links through to its machine-readable definition on apis.io.

Kin Score

This is the API Evangelist rating — a single, repeatable read computed from the artifacts on this page. Green fill is points earned; the red track is points possible, so every bar shows earned-versus-possible at a glance.

Kin Score Kin Score How this is scored →
scored 2026-07-27 · rubric v0.5
Composite quality — 54.6/100 · developing
Contract Quality 14.4 / 25
Developer Ergonomics 9.6 / 20
Commercial Clarity 15.8 / 20
Operational Transparency 6.8 / 13
Governance 0.0 / 12
Discoverability 8.0 / 10
Agent readiness — 48/100 · agent ready
Machine-Readable Contract 18 / 18
Agentic Access Contract 15 / 15
MCP Server 0 / 12
Machine-Readable Auth 10 / 10
Idempotency 0 / 9
Stable Error Semantics 0 / 8
Request/Response Examples 0 / 7
Rate-Limit Signaling 7 / 7
Typed Event Surface 0 / 6
Agent Skills 0 / 5
Well-Known Catalog 0 / 4
Consent & Bot Identity 0 / 3

How we profile LangWatch

Each block below is one kind of artifact we hold for LangWatch. For each we say what it is and why it earns a place in the profile, then list every one we've indexed — capped at two rows, scroll within the panel for the rest.

APIs 28

Each API is captured as its own OpenAPI definition — every operation, parameter, and response. This is the single most useful machine-readable description of what an API does, and it's what lets us score, lint, mock, and generate against it without asking the provider for anything.

Individual APIs this provider publishes, each with its own machine-readable definition.

LangWatch Agents API

The Agents API from LangWatch — 2 operation(s) for agents.

LangWatch Analytics API

The Analytics API from LangWatch — 1 operation(s) for analytics.

LangWatch Annotations API

The Annotations API from LangWatch — 3 operation(s) for annotations.

LangWatch Api Keys API

The Api Keys API from LangWatch — 2 operation(s) for api keys.

LangWatch Budgets API

The Budgets API from LangWatch — 2 operation(s) for budgets.

LangWatch Cache Rules API

The Cache Rules API from LangWatch — 2 operation(s) for cache rules.

LangWatch Dashboards API

The Dashboards API from LangWatch — 3 operation(s) for dashboards.

LangWatch Dataset API

The Dataset API from LangWatch — 7 operation(s) for dataset.

LangWatch Evaluations V3 API

The Evaluations V3 API from LangWatch — 2 operation(s) for evaluations v3.

LangWatch Evaluators API

The Evaluators API from LangWatch — 3 operation(s) for evaluators.

LangWatch Graphs API

The Graphs API from LangWatch — 2 operation(s) for graphs.

LangWatch LangWatch API API

The LangWatch API API from LangWatch — 1 operation(s) for langwatch api.

LangWatch Model Defaults API

The Model Defaults API from LangWatch — 2 operation(s) for model defaults.

LangWatch Model Providers API

The Model Providers API from LangWatch — 2 operation(s) for model providers.

LangWatch Monitors API

The Monitors API from LangWatch — 3 operation(s) for monitors.

LangWatch Projects API

The Projects API from LangWatch — 2 operation(s) for projects.

LangWatch Prompts API

The Prompts API from LangWatch — 8 operation(s) for prompts.

LangWatch Providers API

The Providers API from LangWatch — 2 operation(s) for providers.

LangWatch Scenario Events API

The Scenario Events API from LangWatch — 1 operation(s) for scenario events.

LangWatch Scenarios API

The Scenarios API from LangWatch — 2 operation(s) for scenarios.

LangWatch Secrets API

The Secrets API from LangWatch — 2 operation(s) for secrets.

LangWatch Simulation Runs API

The Simulation Runs API from LangWatch — 3 operation(s) for simulation runs.

LangWatch Suites API

The Suites API from LangWatch — 4 operation(s) for suites.

LangWatch Trace API

The Trace API from LangWatch — 3 operation(s) for trace.

LangWatch Traces API

The Traces API from LangWatch — 3 operation(s) for traces.

LangWatch Triggers API

The Triggers API from LangWatch — 2 operation(s) for triggers.

LangWatch Virtual Keys API

The Virtual Keys API from LangWatch — 4 operation(s) for virtual keys.

LangWatch Workflows API

The Workflows API from LangWatch — 2 operation(s) for workflows.

Scroll within the panel for all 28 ·

Open Collections 1

Open, tool-agnostic collections carry the same runnable value as Postman without locking you to one client — the portable, forkable form of the same exercise.

Open, tool-agnostic API collections (OpenAPI-derived and Bruno).

LangWatch API

OPEN COLLECTION

Pricing Plans 1

Pricing is part of the interface. Machine-readable plans tell you what a tier costs and includes before you commit — one of the six things the Kin Score reads for commercial clarity.

Published pricing tiers and plan structures.

Rate Limits 1

Rate limits are the difference between a demo that works and a production integration that doesn't fall over. Publishing them is an operational-transparency signal — and a hard requirement for any agent that plans its own throughput.

Documented rate limits and quota policies.

Langwatch Rate Limits

4 limits

RATE LIMITS

FinOps 1

Cost, billing, and metering signals let a buyer model the financial operations of an API before it's live. We profile them for the same reason we profile pricing: the money is part of the contract.

Cost, billing, and metering signals for API financial operations.

Features 17

The notable capabilities this provider advertises, captured as structured features so they can be searched and compared instead of read one landing page at a time.

Notable capabilities this provider offers.

OpenTelemetry-native trace ingestion via Python and TypeScript SDKs and any OTel-compliant client
Built-in evaluators — RAGAS, Azure Content Safety, OpenAI Moderation, PII, semantic similarity, language detection, LLM-as-Judge
Real-time online monitors that score production traces against configured evaluators
Batch experiments, suites, and DSPy-driven optimization workflows
Prompt versioning with feature-flag-style deployment and `prompts/sync`/`restore`
Multi-turn agent simulations with User Simulator and Judge Agents (open-source `scenario` framework)
Collaborative annotation and labeling for domain experts and PMs
Datasets — CSV upload, programmatic CRUD, trace-to-record conversion
AI Gateway with OpenAI/Anthropic-compatible proxy, virtual keys, budgets, and semantic cache rules
Open-source Apache-2.0 core, MIT-licensed SDKs, commercial `ee/` enterprise modules
Self-hostable via Docker Compose, Helm, Kind, or full Kubernetes (PostgreSQL + Redis + ClickHouse + OpenSearch)
LangWatch Cloud with three plans — Developer (free), Growth (EUR 59 / core-seat / month), Enterprise (custom)
14-day data retention on free tier, 30-day default on Growth, configurable on Enterprise
MCP server (`@langwatch/mcp-server`) exposing observability, prompts, datasets, scenarios, and evaluator tools to Claude / Cursor / other MCP clients
Integrations — OpenAI, Anthropic, Azure, AWS Bedrock, LiteLLM, LangChain, LangGraph, DSPy, OpenAI Agents, Pydantic AI, CrewAI, Autogen, Haystack
`better-agents` standards repository for agent project scaffolding
SOC 2 / ISO 27001 reports and forward-deployed engineering available under Enterprise tier

Scroll within the panel for all 17 ·

Semantic Vocabularies 1

JSON-LD contexts give the data shared meaning across APIs. We profile them because semantics are what let a machine reconcile 'customer' here with 'customer' somewhere else.

JSON-LD contexts and semantic vocabularies used across these APIs.

Langwatch Context

56 classes · 0 properties

JSON-LD

Security Posture 3

Authentication, domain security, vulnerability disclosure, and trust-center signals — the evidence that a provider takes security seriously enough to document it. We profile it because you can't govern what you can't see.

Authentication, domain security, vulnerability disclosure, and trust-center signals.

Langwatch Authentication

apiKey/http · 2 schemes

SECURITY

Langwatch Domain Security

TLSv1.3 · HSTS · DMARC

SECURITY

Langwatch Trust Center

ISO 27001, GDPR

SECURITY

Agentic Access 1

An x-agentic-access contract marks which operations are safe for an agent to run on its own and which need a human in the loop. It is the difference between an API an agent can use and one it can use safely.

Recommended x-agentic-access execution contracts for AI agents.

Langwatch Agentic Access

133 operations · 85 acting · 3 human-in-the-loop

133 operations · 85 acting

AGENTIC

Resources

Every other property we hold for LangWatch — documentation, portals, status pages, policies, and corporate surface — grouped by the job it does, following the integrator's arc from getting started to running in production.

Get Started 3

Portal, sign-up, and the first successful call

Agent Surfaces 2

MCP servers, agent skills, and machine-readable catalogs

Access & Security 3

Authentication, authorization, and security posture

Learn 1

Tutorials, courses, talks, and written guidance

Operate 3

Status, limits, changes, and where to get help

Company 3

The organization behind the API

Other 1

Properties that don't map to a standard resource type

← All providers · Data indexed from github.com/api-evangelist/langwatch · machine-readable index on apis.io