Galileo
Evaluation, observability and guardrails for LLM applications and agents, built for enterprise scale rather than for a prototype.
Highlights
- Trace capture and observability for GenAI apps and multi-agent systems
- Research-backed evaluation metrics for hallucination, correctness, and quality
- Luna and Luna-2 evaluation models for fast, low-cost scoring at scale
- Offline evals and experiments to compare prompts, models, and agent versions
- Real-time guardrails that block hallucinations and unsafe outputs in production
- Agentic evaluations purpose-built for tool calls and multi-step agent traces
- Production monitoring with alerting on quality and reliability regressions
- Python and TypeScript SDKs plus enterprise deployment options
External link — opens galileo.ai in a new tab. Galileo is a third-party product; we are not affiliated with it.
About Galileo
What it is
Galileo tests, monitors and guardrails generative AI applications and agents — running evaluations before release, tracing behaviour in production, and blocking output at runtime. It is aimed at enterprises operating AI at scale. Not to be confused with the text-to-UI design tool of the same name.
Why it's different
It competes with LangSmith, Langfuse, Freeplay and Future AGI, and the field is crowded enough that feature lists do not separate them. Galileo's positioning is enterprise depth — scale, access control, the reporting an organisation needs to answer for a model's behaviour — rather than developer ergonomics. That is a real split: a team wanting to wire up tracing in an afternoon will find the open-source options faster, and a bank needing to demonstrate governance will not find them sufficient. There is no meaningful self-hosted free path here, which decides it for some teams before any evaluation.
How people use it
It is used by organisations with AI in front of customers, where a bad output is an incident rather than a bug. The workflow that actually pays is domain experts labelling real production traces, those labels becoming the evaluation set, and every change running against it. That loop is the whole value and it fails in exactly one way, which is nobody doing the labelling.
Written by the n3os team. We are not affiliated with Galileo.
This listing was written from public information, without Galileo’s involvement. If you own it and something here is wrong — or you would rather not be listed at all — email us and we will correct or remove it.
Get the ones worth knowing about
We write one of these for every tool worth the trouble. Get the new ones, plus what we have found genuinely useful lately.
Your address goes to Buttondown, who send the emails on our behalf. One click unsubscribes, and the list is never sold or shared.