Need help with your APIs? I offer API discovery, governance & evangelism services. Explore services →
API Evangelist API Evangelist
Discovery
Learnings
Guidance
Toolbox
Alignment
API Evangelist LLC

Promptfoo

Promptfoo is an open-source CLI and TypeScript/Node.js library for evaluating, red-teaming, and security-testing LLM applications, agents, and RAG pipelines. It runs deterministic prompt evals with model-graded and rule-based assertions, generates dynamic adversarial attack probes across 50+ vulnerability categories (prompt injection, jailbreaks, RAG poisoning, PII leakage, harmful content, business rule violations), and integrates with CI/CD via a GitHub Action and a `code-scans` command for pull-request review. Promptfoo supports dozens of LLM providers — OpenAI, Anthropic Claude, Google Gemini, AWS Bedrock/SageMaker, Azure OpenAI, Mistral, Cohere, Groq, DeepSeek, Together, Fireworks, OpenRouter, LiteLLM, Vercel and Cloudflare AI gateways — plus local runtimes (Ollama, LocalAI, llama.cpp, vLLM, llamafile, Docker Model Runner) and custom HTTP/WebSocket/Python/JavaScript/Go/Ruby/Shell providers. It also ships an MCP server (`promptfoo mcp`), a Model Audit scanner (`scan-model`) for malicious ML artifacts, and a hosted Enterprise tier with team sharing, continuous monitoring, SSO, and a centralized compliance dashboard aligned to OWASP LLM Top 10, NIST AI RMF, MITRE ATLAS, and the EU AI Act. The project is MIT-licensed, has 21k+ GitHub stars, and is now part of OpenAI while remaining open source.

human only

More than an index entry, but the surface is still mostly links rather than artifacts — the cohort most likely to move a full band from modest, well-targeted work.

Kin Score

API Evangelist profiles Promptfoo the way a machine reads it — 26 machine-readable artifacts, pulled from the provider's own public surface and indexed so a developer, an analyst, or an AI agent can evaluate it against every other provider on the network.

Every provider in the network is reduced to the same set of machine-readable artifacts — OpenAPI contracts, event specifications, GraphQL schemas, runnable collections, pricing and rate-limit signals, security posture, OAuth scopes, and the agent surfaces (MCP servers and skills) that let software drive the API on its own. We profile them because the interface is the part of a company you can actually inspect: it is a truer signal of what a provider does than any marketing page. From those artifacts we compute the Kin Score — Promptfoo scores 22.6/100 (emerging), with a separate agent-readiness read of 7/100 (human only). The full breakdown is below, followed by every artifact we hold — each card links through to its machine-readable definition on apis.io.

Kin Score

This is the API Evangelist rating — a single, repeatable read computed from the artifacts on this page. Green fill is points earned; the red track is points possible, so every bar shows earned-versus-possible at a glance.

Kin Score Kin Score How this is scored →
scored 2026-07-27 · rubric v0.5
Composite quality — 22.6/100 · emerging
Contract Quality 0.0 / 25
Developer Ergonomics 7.4 / 20
Commercial Clarity 3.7 / 20
Operational Transparency 4.8 / 13
Governance 0.0 / 12
Discoverability 6.8 / 10
Agent readiness — 7/100 · human only
Machine-Readable Contract 0 / 18
Agentic Access Contract 0 / 15
MCP Server 0 / 12
Machine-Readable Auth 0 / 10
Idempotency 0 / 9
Stable Error Semantics 0 / 8
Request/Response Examples 7 / 7
Rate-Limit Signaling 0 / 7
Typed Event Surface 0 / 6
Agent Skills 0 / 5
Well-Known Catalog 0 / 4
Consent & Bot Identity 0 / 3

How we profile Promptfoo

Each block below is one kind of artifact we hold for Promptfoo. For each we say what it is and why it earns a place in the profile, then list every one we've indexed — capped at two rows, scroll within the panel for the rest.

Features 24

The notable capabilities this provider advertises, captured as structured features so they can be searched and compared instead of read one landing page at a time.

Notable capabilities this provider offers.

Open-source CLI and Node.js/TypeScript library for LLM evaluation and red teaming
`promptfoo eval` — run deterministic prompt/model evals with caching, concurrency, and live reload
`promptfoo redteam` — generate and run dynamic adversarial probes across 50+ vulnerability categories
`promptfoo view` — local browser UI for browsing eval and red-team results
`promptfoo share` — publish a shareable URL for an eval or model audit
`promptfoo generate` — synthesize datasets, red-team tests, and assertions
`promptfoo optimize` — improve prompts against a target provider
`promptfoo scan-model` — security-scan ML model files (ModelAudit)
`promptfoo code-scans` — scan code changes for LLM security vulnerabilities in IDEs and CI/CD
`promptfoo mcp` — expose promptfoo tools as a Model Context Protocol server
`promptfoo retry`, `list`, `export`, `import`, `validate`, `debug`, `cache`, `auth`
Assertions: rule-based (equals, contains, regex, javascript, python, cost, latency) and model-graded (llm-rubric, classifier, factuality, answer-relevance, similarity)
Red-team plugins covering prompt injection, jailbreaks, PII leakage, harmful content, bias, business-rule violations, RAG poisoning, and agent/tool abuse
Attack strategies including multi-turn (Crescendo), GOAT (Meta), and iterative jailbreak techniques
Framework alignment: OWASP LLM Top 10, NIST AI RMF, MITRE ATLAS, EU AI Act
50+ providers: OpenAI, Anthropic, Google Gemini, AWS Bedrock, SageMaker, Azure OpenAI, Mistral, Cohere, Groq, DeepSeek, Together, Fireworks, Perplexity, OpenRouter, LiteLLM, Vercel AI Gateway, Cloudflare AI Gateway
Local runtimes: Ollama, LocalAI, llama.cpp, vLLM, Docker Model Runner, llamafile
Custom providers via HTTP, WebSocket, Python, JavaScript, Go, Ruby, and Shell
GitHub Action (`promptfoo/promptfoo-action`) for PR-level eval gating
Self-hostable via Helm chart and on-premise enterprise deployment
Hosted Enterprise tier: team sharing, continuous monitoring, centralized compliance dashboard, SSO, granular permissions, managed cloud, SLA-backed support
SOC 2 and ISO 27001 certified; trust center at trust.promptfoo.dev
Distributed on npm (`promptfoo`), PyPI (`promptfoo`), and Homebrew (`brew install promptfoo`)
MIT-licensed; 21k+ GitHub stars; now part of OpenAI

Scroll within the panel for all 24 ·

Security Posture 2

Authentication, domain security, vulnerability disclosure, and trust-center signals — the evidence that a provider takes security seriously enough to document it. We profile it because you can't govern what you can't see.

Authentication, domain security, vulnerability disclosure, and trust-center signals.

Prompt Foo Domain Security

TLSv1.3 · DMARC

SECURITY

Prompt Foo Trust Center

SOC 2, ISO 27001

SECURITY

Resources

Every other property we hold for Promptfoo — documentation, portals, status pages, policies, and corporate surface — grouped by the job it does, following the integrator's arc from getting started to running in production.

Get Started 2

Portal, sign-up, and the first successful call

Access & Security 3

Authentication, authorization, and security posture

Operate 3

Status, limits, changes, and where to get help

Commercial 2

Pricing, plans, and the legal terms of use

Other 1

Properties that don't map to a standard resource type

← All providers · Data indexed from github.com/api-evangelist/prompt-foo · machine-readable index on apis.io