监控、日志、追踪、性能分析
langfuse
🪢 Open source LLM engineering platform: LLM Observability, metrics, evals, prompt management, playground, datasets. Integrates with OpenTelemetry, Langchain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
confident-ai
DeepEval is an open-source LLM evaluation framework that enables unit testing for LLMs and RAG pipelines, featuring 14+ evaluation metrics, synthetic test data generation, and CI/CD integration for continuous quality assurance.
Arize-ai
Phoenix is an open-source AI observability and evaluation platform by Arize AI, providing LLM tracing, prompt engineering tools, evaluation datasets, and visualization for debugging and improving AI applications in production.
Helicone
🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 🍓
AgentOps-AI
Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks including CrewAI, Agno, OpenAI Agents SDK, Langchain, Autogen, AG2, and Ca