How AI is applied across API Evangelist and APIs.io. Read my AI disclosure →
API Evangelist API Evangelist
Discovery
Learnings
Guidance
Toolbox
Alignment
API Evangelist LLC
LangWatch website screenshot

LangWatch

LangWatch is an open-source LLM observability, evaluation, and AI agent testing platform. Built around OpenTelemetry-native tracing, LangWatch lets teams instrument LLM applications (LangChain, LangGraph, DSPy, OpenAI Agents, LiteLLM, Pydantic AI, CrewAI, AWS Bedrock, and more), run real-time and batch evaluations, version and deploy prompts, simulate multi-turn agent conversations against scripted scenarios and Judge Agents, manage labeled datasets, and govern model traffic through a virtual-key AI Gateway with budgets and semantic caching. The platform exposes a REST API at app.langwatch.ai, ships Python and TypeScript SDKs plus an MCP server, publishes the companion `scenario` agent-testing framework and `better-agents` standards, and runs on Apache-2.0 core (with `ee/` enterprise modules under commercial license) — deployable as LangWatch Cloud or self-hosted via Docker Compose, Helm, Kind, or full Kubernetes.

agent ready

Real signal across most facets with visible, nameable gaps — the contract exists but is thin, or the portal is good while governance and commercial terms are absent.

Kin Score

API Evangelist profiles LangWatch the way a machine reads it — 83 machine-readable artifacts across 1 API, pulled from the provider's own public surface and indexed so a developer, an analyst, or an AI agent can evaluate it against every other provider on the network.

Every provider in the network is reduced to the same set of machine-readable artifacts — OpenAPI contracts, event specifications, GraphQL schemas, runnable collections, pricing and rate-limit signals, security posture, OAuth scopes, and the agent surfaces (MCP servers and skills) that let software drive the API on its own. We profile them because the interface is the part of a company you can actually inspect: it is a truer signal of what a provider does than any marketing page. From those artifacts we compute the Kin Score — LangWatch scores 47.1/100 (developing), with a separate agent-readiness read of 36/100 (agent ready). The full breakdown is below, followed by every artifact we hold — each card links through to its machine-readable definition on apis.io.

Kin Score

This is the API Evangelist rating — a single, repeatable read computed from the artifacts on this page. Green fill is points earned; the red track is points possible, so every bar shows earned-versus-possible at a glance. Every facet and dimension name is a link: it opens that measurement's page on APIs.io, where the rating runs across the whole catalog — the exact checks that feed it, how every profiled provider distributes on it, and who is at the top of it.

Kin Score Kin Score How this is scored →
scored 2026-09-08 · rubric v0.20.0
Composite quality — 47.1/100 · developing
Contract Quality 14.8 / 25
Access Clarity 12.1 / 20
Discoverability 5.4 / 10
LangWatch Kin Score — API readiness rating by API Evangelist

Put this on your own site. The badge is drawn live from LangWatch's current Kin Score — paste it once and it updates itself every time the score is recomputed. It follows your visitor's light or dark setting, and it links back here so anyone who sees it can read the full breakdown.

<!-- Kin Score · API Evangelist -->
<a href="https://providers.apievangelist.com/providers/langwatch/"
   title="LangWatch on API Evangelist — API profile and Kin Score">
  <img src="https://apis.io/badge/langwatch.svg"
       alt="LangWatch Kin Score — API readiness rating by API Evangelist" width="150" height="150" loading="lazy">
</a>

More shapes, themes and sizes → · Score as JSON · How badges work

How we profile LangWatch

Each block below is one kind of artifact we hold for LangWatch. For each we say what it is and why it earns a place in the profile, then list every one we've indexed — capped at two rows, scroll within the panel for the rest.

APIs 28

Each API is captured as its own OpenAPI definition — every operation, parameter, and response. This is the single most useful machine-readable description of what an API does, and it's what lets us score, lint, mock, and generate against it without asking the provider for anything.

Individual APIs this provider publishes, each with its own machine-readable definition.

LangWatch Agents API

The Agents API from LangWatch — 2 operation(s) for agents.

LangWatch Analytics API

The Analytics API from LangWatch — 1 operation(s) for analytics.

LangWatch Annotations API

The Annotations API from LangWatch — 3 operation(s) for annotations.

LangWatch Api Keys API

The Api Keys API from LangWatch — 2 operation(s) for api keys.

LangWatch Budgets API

The Budgets API from LangWatch — 2 operation(s) for budgets.

LangWatch Cache Rules API

The Cache Rules API from LangWatch — 2 operation(s) for cache rules.

LangWatch Dashboards API

The Dashboards API from LangWatch — 3 operation(s) for dashboards.

LangWatch Dataset API

The Dataset API from LangWatch — 7 operation(s) for dataset.

LangWatch Evaluations V3 API

The Evaluations V3 API from LangWatch — 2 operation(s) for evaluations v3.

LangWatch Evaluators API

The Evaluators API from LangWatch — 3 operation(s) for evaluators.

LangWatch Graphs API

The Graphs API from LangWatch — 2 operation(s) for graphs.

LangWatch LangWatch API API

The LangWatch API API from LangWatch — 1 operation(s) for langwatch api.

LangWatch Model Defaults API

The Model Defaults API from LangWatch — 2 operation(s) for model defaults.

LangWatch Model Providers API

The Model Providers API from LangWatch — 2 operation(s) for model providers.

LangWatch Monitors API

The Monitors API from LangWatch — 3 operation(s) for monitors.

LangWatch Projects API

The Projects API from LangWatch — 2 operation(s) for projects.

LangWatch Prompts API

The Prompts API from LangWatch — 8 operation(s) for prompts.

LangWatch Providers API

The Providers API from LangWatch — 2 operation(s) for providers.

LangWatch Scenario Events API

The Scenario Events API from LangWatch — 1 operation(s) for scenario events.

LangWatch Scenarios API

The Scenarios API from LangWatch — 2 operation(s) for scenarios.

LangWatch Secrets API

The Secrets API from LangWatch — 2 operation(s) for secrets.

LangWatch Simulation Runs API

The Simulation Runs API from LangWatch — 3 operation(s) for simulation runs.

LangWatch Suites API

The Suites API from LangWatch — 4 operation(s) for suites.

LangWatch Trace API

The Trace API from LangWatch — 3 operation(s) for trace.

LangWatch Traces API

The Traces API from LangWatch — 3 operation(s) for traces.

LangWatch Triggers API

The Triggers API from LangWatch — 2 operation(s) for triggers.

LangWatch Virtual Keys API

The Virtual Keys API from LangWatch — 4 operation(s) for virtual keys.

LangWatch Workflows API

The Workflows API from LangWatch — 2 operation(s) for workflows.

Scroll within the panel for all 28 ·

Open Collections 30

Open, tool-agnostic collections carry the same runnable value as Postman without locking you to one client — the portable, forkable form of the same exercise.

Open, tool-agnostic API collections (OpenAPI-derived and Bruno).

Scroll within the panel for all 30 ·

Pricing Plans 1

Pricing is part of the interface. Machine-readable plans tell you what a tier costs and includes before you commit — one of the things the Kin Score reads for access clarity — renamed from commercial clarity in rubric 0.12, because a free statutory interface has access terms and no commercial ones.

Published pricing tiers and plan structures.

Rate Limits 1

Rate limits are the difference between a demo that works and a production integration that doesn't fall over. Publishing them is an operational-transparency signal — and a hard requirement for any agent that plans its own throughput.

Documented rate limits and quota policies.

Langwatch Rate Limits

4 limits

RATE LIMITS

FinOps 1

Cost, billing, and metering signals let a buyer model the financial operations of an API before it's live. We profile them for the same reason we profile pricing: the money is part of the contract.

Cost, billing, and metering signals for API financial operations.

Features 17

The notable capabilities this provider advertises, captured as structured features so they can be searched and compared instead of read one landing page at a time.

Notable capabilities this provider offers.

OpenTelemetry-native trace ingestion via Python and TypeScript SDKs and any OTel-compliant client
Built-in evaluators — RAGAS, Azure Content Safety, OpenAI Moderation, PII, semantic similarity, language detection, LLM-as-Judge
Real-time online monitors that score production traces against configured evaluators
Batch experiments, suites, and DSPy-driven optimization workflows
Prompt versioning with feature-flag-style deployment and `prompts/sync`/`restore`
Multi-turn agent simulations with User Simulator and Judge Agents (open-source `scenario` framework)
Collaborative annotation and labeling for domain experts and PMs
Datasets — CSV upload, programmatic CRUD, trace-to-record conversion
AI Gateway with OpenAI/Anthropic-compatible proxy, virtual keys, budgets, and semantic cache rules
Open-source Apache-2.0 core, MIT-licensed SDKs, commercial `ee/` enterprise modules
Self-hostable via Docker Compose, Helm, Kind, or full Kubernetes (PostgreSQL + Redis + ClickHouse + OpenSearch)
LangWatch Cloud with three plans — Developer (free), Growth (EUR 59 / core-seat / month), Enterprise (custom)
14-day data retention on free tier, 30-day default on Growth, configurable on Enterprise
MCP server (`@langwatch/mcp-server`) exposing observability, prompts, datasets, scenarios, and evaluator tools to Claude / Cursor / other MCP clients
Integrations — OpenAI, Anthropic, Azure, AWS Bedrock, LiteLLM, LangChain, LangGraph, DSPy, OpenAI Agents, Pydantic AI, CrewAI, Autogen, Haystack
`better-agents` standards repository for agent project scaffolding
SOC 2 / ISO 27001 reports and forward-deployed engineering available under Enterprise tier

Scroll within the panel for all 17 ·

Semantic Vocabularies 1

JSON-LD contexts give the data shared meaning across APIs. We profile them because semantics are what let a machine reconcile 'customer' here with 'customer' somewhere else.

JSON-LD contexts and semantic vocabularies used across these APIs.

Langwatch Context

56 classes · 0 properties

JSON-LD

Security Posture 3

Authentication, domain security, vulnerability disclosure, and trust-center signals — the evidence that a provider takes security seriously enough to document it. We profile it because you can't govern what you can't see.

Authentication, domain security, vulnerability disclosure, and trust-center signals.

Langwatch Authentication

apiKey/http · 2 schemes

SECURITY

Langwatch Domain Security

TLSv1.3 · HSTS · DMARC

SECURITY

Langwatch Trust Center

ISO 27001, GDPR

SECURITY

Agentic Access 1

An x-agentic-access contract marks which operations are safe for an agent to run on its own and which need a human in the loop. It is the difference between an API an agent can use and one it can use safely.

Recommended x-agentic-access execution contracts for AI agents.

Langwatch Agentic Access

133 operations · 85 acting · 3 human-in-the-loop

133 operations · 85 acting

AGENTIC

Resources

Every other property we hold for LangWatch — documentation, portals, status pages, policies, and corporate surface — grouped by the job it does, following the integrator's arc from getting started to running in production.

Get Started 3

Portal, sign-up, and the first successful call

Agent Surfaces 1

MCP servers, agent skills, and machine-readable catalogs

Learn 1

Tutorials, courses, talks, and written guidance

Other 2

Properties that don't map to a standard resource type

← All providers · Data indexed from github.com/api-evangelist/langwatch · machine-readable index on apis.io

Where this information came from

This is an independent, third-party profile of LangWatch, published by API Evangelist. We do not operate, host, resell, or support these APIs, and we are not affiliated with or endorsed by the company unless stated above. Everything here is built from publicly available information — the company's own site, developer portal, documentation, public repositories, and the specifications it publishes for public use. Nothing is obtained by breaching a system, defeating an access control, or using credentials.

The Kin Score and Agent Readiness rating are independently calculated assessments of a company's public API artifacts, scored against a published rubric. They are not certifications, endorsements, security assessments, or audits.

Corrections, re-scores, and removal are free — no partnership or purchase required, and you do not need to justify the request. A removed company is recorded as unrated, never scored zero for having asked. Acknowledgement within one business day; removal within two.

info@apievangelist.com · Read the full data-sourcing policy →
On a security or compliance team? Put security in the subject line and you will get a person, not a form — we will tell you exactly which public URLs this profile was built from.