tools
Promptfoo
Promptfoo evaluates and red-teams LLM apps, prompts, models, agents, and RAG systems through a CLI, library, and web interface.

In inglese
Promptfoo helps developers test prompt and model quality, compare outputs, define metrics, and run evaluations in CI/CD. It also scans AI applications for vulnerabilities through automated red teaming, pentesting, and code scanning.
It runs locally or can be self-hosted, supports custom LLM APIs, and provides team sharing and enterprise monitoring. The Community plan is free; Enterprise and On-Premise plans are custom-priced, and additional red-team probes may cost extra.
Features
- Evaluate prompts, models, agents, and RAG pipelines
- Run automated red teaming and vulnerability scanning
- Compare multiple prompts and providers side by side
- Define automatic metrics and assertions
- Use caching, concurrency, and live reloading
- Integrate evaluations into CI/CD workflows
- Run as a CLI, library, or local web viewer
- Open-source under the MIT license
Use cases
- Benchmark prompts and models for a specific application
- Scan an AI agent for prompt injection and other vulnerabilities
- Compare model responses before changing a production provider
- Run LLM quality and security checks in CI/CD
- Test RAG systems for accuracy and source attribution
- Review LLM-related security issues in pull requests
Pros
Cons
Latest updates
- Open-Sourcing ModelAudit: Security Scanner for ML Model Files
ModelAudit was released as an MIT-licensed open-source static scanner for ML model files.
- Building a Security Scanner for LLM Apps
Promptfoo is adding code scanning for LLM-related vulnerabilities, first as a GitHub Action that reviews pull requests.
- GPT-5.2 Initial Trust and Safety Assessment (GPT-5.2)
Promptfoo opened a PR for GPT-5.2 support.
- Real-Time Fact Checking for LLM Outputs
Promptfoo added a search-rubric assertion that uses a separate judge to verify answers against current information.
Capabilities
- Command line — “promptfoo is an open-source CLI and library for evaluating and red-teaming LLM apps.” source
- Choice of models — “Use OpenAI, Anthropic, Azure, Google, HuggingFace, open-source models like Llama, or integrate custom API providers for any LLM API” source
- Self-hosted — “Run locally or self-host on your own infrastructure” source
- API — “Promptfoo API access” source
- Traces and evaluates — “Test and evaluate your prompts, models, and RAG pipelines” source
Get it
Security
Pricing
- Starting price
- Custom
- Prices checked
- 2026-09-24
Community
Free Forever
- All LLM evaluation features
- All model providers and integrations
- Red teaming (10k probes/month)
- Custom integration with your own app
- Run locally or self-host on your own infrastructure
- Vulnerability scanning
- Community support
Enterprise
Custom
- Custom red teaming limits
- Team sharing & collaboration
- Continuous monitoring
- Centralized security/compliance dashboard
- SSO and granular permission profiles
- Promptfoo API access
- Managed cloud deployment
- Priority support & SLA guarantees
On-Premise
Custom
- All Enterprise features
- Deployment on your own infrastructure
- Complete data isolation
- Dedicated runner
- Assigned deployment engineer