DeepEval
@deepeval.com
DeepEval is the open-source LLM evaluation framework for testing and benchmarking LLM applications — 50+ plug-and-play metrics for AI agents, RAG, chatbots, and more.
DeepEval's Company Logos
DeepEval's Brand Colors
Hex Code
Color name
RGB
HSL
CMYK
#6200FF
Electric Violet
98, 0, 255
263, 100, 50
62, 100, 0, 0
#034EA2
Endeavour
3, 78, 162
212, 96, 32
98, 52, 0, 36
#FBFAF6
Spring Wood
251, 250, 246
48, 38, 97
0, 0, 2, 2
About DeepEval
DeepEval is an open-source framework for evaluating and testing large language model applications, including AI agents, retrieval-augmented generation systems, and chatbots. Used by developers and leading AI companies, it helps teams build reliable evaluation pipelines and catch regressions as they develop and deploy AI systems.
The framework offers more than 50 research-backed metrics covering areas such as hallucination, faithfulness, answer relevance, summarization, toxicity, and bias. It also supports conversational, voice, and multimodal evaluation across text, images, and audio. Developers can run tests in Python or integrate them into CI/CD workflows, inspect execution traces, and assess individual steps with scored, explainable results.
DeepEval includes tools for generating synthetic test cases from knowledge bases and simulating conversations, helping teams evaluate applications before launch. Its evaluation techniques include criteria-based scoring, decision-graph metrics, and reference-grounded question answering. The framework integrates with a wide range of models, agent frameworks, and development pipelines. For teams seeking broader collaboration and operational capabilities, DeepEval also connects with Confident AI, a platform offering observability, production monitoring, dataset management, and related evaluation features.
Brand industry
Computers Electronics and Technology
Company type
Suggest company type
Year founded
Suggest founded year
Company size
Suggest company size
