Galileo AI
FreemiumEvaluation intelligence platform for generative AI and agentic applications.
Galileo AI is an evaluation intelligence platform for generative AI and agentic applications. It combines observability, experimentation, and runtime protection so AI teams can evaluate, iterate, monitor, and protect LLM-powered systems at scale. The platform captures sessions, traces, and spans from applications; provides out-of-the-box metrics across response quality, RAG, agentic performance, safety, and compliance; and supports custom LLM-as-a-judge and code-based metrics. A free tier includes 5,000 traces per month, with paid plans for higher volumes and enterprise features.

What is it
An evaluation intelligence platform that helps teams observe, evaluate, experiment on, and protect generative AI and agentic applications.
What it can do
Log traces and spans, run experiments with datasets, apply out-of-the-box and custom metrics, detect hallucinations and safety risks, and enforce runtime guardrails through centralized controls.
Who is it for
AI engineers, product managers, ML teams, and enterprises that need systematic evaluation and production guardrails for LLM and agent systems.
Key Features
AI Observability – Session, Trace, and Span Logging
Captures structured runtime data from GenAI applications, organizing it into projects, log streams, sessions, traces, and spans for full visibility.
Out-of-the-Box Metrics – Seven Evaluation Categories
Provides ready-to-use metrics for agentic performance, response quality, RAG, safety and compliance, expression and readability, multimodal quality, and text-to-SQL.
LLM-as-a-Judge and Custom Metrics
Lets teams create custom evaluation metrics using LLM judges or custom code, then improve them continuously with feedback-driven Autotune.
Experiments – Systematic Prompt and Model Comparison
Runs datasets against prompts, models, or application code to compare configurations and measure improvements with defined metrics.
Use Cases
Evaluate RAG Pipeline Quality
Measure retrieval accuracy, context relevance, and generation groundedness to ensure RAG systems produce reliable, well-sourced answers.
Monitor LLM Applications in Production
Log traces and spans to understand runtime behavior, latency, token usage, cost, and metric trends over time.
Detect Hallucinations in AI Outputs
Use correctness and ground-truth adherence metrics to identify factual errors and unsupported claims in generated content.
Run Prompt Engineering Experiments
Compare prompts, models, and configurations against datasets to find the highest-performing setup for a given use case.
Pricing plans
Frequently Asked Questions
Discussion
No comments yet. Be the first to start the thread.