SEARWEB SITE INDEX

Braintrust · Indexed content

braintrust.dev

Explore internal pages, articles and content excerpts discovered from this site’s public sources.

Internal links
466
Articles
397
Last indexed
21/9/2026, 4:01:17
466 indexed items
ArticleSitemap

2026 10 07 evals foundations

https://www.braintrust.dev/workshops/2026-10-07-evals-foundations

Open original page

Designed for new Braintrust users, this workshop will cover the core foundations of evals and how to get started with Braintrust.

Language: en
Indexed excerpt

WorkshopsEvals foundations7 October 202610:00 AM PTPMsJess WangDeveloper AdvocateRegister for the workshopDesigned for new Braintrust users, this workshop will cover the core foundations of evals and how to get started with Braintrust. Jess will walk through the practical pieces of the eval workflow: moving from a question about your AI system to experiments, results, and feedback you can use to improve it. Bring a system you are building, a quality question you want to answer, or simply your curiosity. You will leave with a clearer picture of how to measure and improve agent quality, plus a practical next step to take. What you’ll learn How Braintrust groups production traces into named failure patterns you can act on How to filter a failure cluster into a labeled dataset How to write an eval that targets a specific failure pattern and validate the fix How to run a repeatable diagnose-to-eval workflow in Braintrust Recent on-demand workshopsEvals foundations9 September 202610:00 AM PTDesigned for new Braintrust users, this workshop will cover the core foundations of evals and how to get started with Braintrust.PMsHow financial services teams ship AI agents to production26 August 2

Discovered: Last checked: Content changed:
ArticleSitemap

Non-determinism

https://www.braintrust.dev/encyclopedia/non-determinism

Open original page

The property of agents where the same input can produce different outputs across identical requests.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Multimodal

https://www.braintrust.dev/encyclopedia/multimodal

Open original page

Agents or data that involve more than one type of input or output, such as text, images, audio, or PDFs. Multimodal traces and datasets require observability infrastructure that can store and display non-text content alongside structured data.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Model drift

https://www.braintrust.dev/encyclopedia/model-drift

Open original page

A change in model behavior or performance over time, often caused by shifts in input distribution or provider updates. It's typically detected via score trends, topic shifts, or regressions on a fixed eval suite.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Multimodal dataset

https://www.braintrust.dev/encyclopedia/multimodal-dataset

Open original page

A dataset containing non-text inputs/outputs (images, audio, PDFs), requiring multimodal evals and storage. Multimodal datasets often need specialized rubrics and display tooling for review.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Model comparison

https://www.braintrust.dev/encyclopedia/model-comparison

Open original page

Evaluating multiple models against the same dataset and scorers to determine which performs best for a given use case. Comparisons are most useful when cost and latency are included, not just quality.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Loop

https://www.braintrust.dev/encyclopedia/loop

Open original page

Braintrust's agent that helps teams write scorers, iterate on prompts, generate datasets, run evals, and surface patterns in production data. Loop reduces the friction of going from "something broke" to "we shipped a measured fix."

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Logs

https://www.braintrust.dev/encyclopedia/logs

Open original page

Capturing structured execution data from both development and production environments. In agent observability, logs are trace-based rather than line-based.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Metadata

https://www.braintrust.dev/encyclopedia/metadata

Open original page

Optional key-value pairs attached to a dataset record or trace, used for filtering, grouping, and analysis. Metadata lets you slice performance by segment (tier, language, intent, etc.).

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Instrumentation

https://www.braintrust.dev/encyclopedia/instrumentation

Open original page

Adding code/hooks to capture traces, metadata, and events throughout the system. Good instrumentation makes failures reproducible and measurable instead of anecdotal.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Inter-annotator agreement (IAA)

https://www.braintrust.dev/encyclopedia/inter-annotator-agreement-iaa

Open original page

A measure of how consistently multiple humans label the same items. Low IAA is usually a sign the rubric or schema needs clearer definitions.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

LLM-as-a-judge

https://www.braintrust.dev/encyclopedia/llm-as-a-judge

Open original page

A scoring approach where a language model evaluates the quality or correctness of another model's output, typically guided by a rubric. It enables scalable evals of open-ended outputs.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Label

https://www.braintrust.dev/encyclopedia/label

Open original page

A tag or categorical score applied to a trace or dataset record during annotation. Labels make it possible to filter, route, and prioritize work.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Hallucination

https://www.braintrust.dev/encyclopedia/hallucination

Open original page

When a model produces output that is not supported by the provided context or by reliable external facts, but presents it as if it were true. Hallucinations are a common form of semantic failure.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Hybrid deployment

https://www.braintrust.dev/encyclopedia/hybrid-deployment

Open original page

A deployment model where the control plane is hosted by the vendor and the data plane runs in the customer's cloud environment. Hybrid deployment keeps customer data in their VPC while reducing operational overhead.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Human review

https://www.braintrust.dev/encyclopedia/human-review

Open original page

A workflow in which team members manually examine and score traces or dataset records through a dedicated UI. Human review is often reserved for high-risk or high-impact cases.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Input

https://www.braintrust.dev/encyclopedia/input

Open original page

The data passed to a task function during an eval. Inputs can be raw text, structured objects, or multimodal attachments.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Gateway

https://www.braintrust.dev/encyclopedia/gateway

Open original page

A unified API layer that routes LLM requests to any supported provider through a single interface. Gateways help provide automatic tracing and consistent telemetry across models.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleInternal link

The principles of Braintrust

https://www.braintrust.dev/manifesto

Open original page

Braintrust's principles, beliefs, and practical guidance for building high-quality evals.

Language: en
Indexed excerpt

Evals and observability are AI infrastructureThe eval is as important as the model, harness, etc. It should be treated as part of the stack, not an afterthought.If you’re building agents, you need to consider the infrastructure stack the same way you would for any software. What cloud provider, database, what IaC framework? What experiments, which scorers, which rubrics?Agents are now in production, supporting products used by millions. We are past the point of toy demos and neat experiments. Serious products need serious infrastructure. And real infrastructure is observable.

Discovered: Last checked: Content changed:
ArticleInternal link

FoundationsLearn to eval

https://www.braintrust.dev/foundations

Open original page

A free course on LLM evals showing how to build, observe, and eval real AI systems.

Language: en
Indexed excerpt

{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://www.braintrust.dev"},{"@type":"ListItem","position":2,"name":"Evals","item":"https://www.braintrust.dev/foundations"}]}ProductResourcesCustomersPricingContact usSign inSign upObserveTrace everythingEvaluateTest what shipsDiscoverFind behaviorsDocsStart buildingBlogInsights and updates .workshops-icon-col { transition: transform 400ms cubic-bezier(0.65, 0, 0.35, 1); } .workshops-icon-playing .workshops-icon-col { transform: translateY(7.5px); } WorkshopsLive sessions .evals-icon-row { transition: transform 350ms ease-in-out; } .evals-icon-playing .evals-icon-row-down { transform: translateY(5px); } .evals-icon-playing .evals-icon-row-up { transform: translateY(-5px); } Eval libraryOpen source researchFoundationsLearn to evalEncyclopediaEval conceptsProductObserveEvaluateDiscoverResourcesDocsBlogWorkshopsEval libraryFoundationsEncyclopediaCustomersPricingContact usSign inSign upFoundationsEvalsLearn how the best teams, including Ramp, Notion, and OpenAI, ship quality AI products. In this course, you'll build a customer support chatbot from scratc

Discovered: Last checked: Content changed:
ArticleInternal link

Take the eval maturity assessment

https://www.braintrust.dev/assessment

Open original page

Find gaps in your AI development loop and get an AI quality roadmap.

Language: en
Indexed excerpt

ProductResourcesCustomersPricingContact usSign inSign upObserveTrace everythingEvaluateTest what shipsDiscoverFind behaviorsDocsStart buildingBlogInsights and updates .workshops-icon-col { transition: transform 400ms cubic-bezier(0.65, 0, 0.35, 1); } .workshops-icon-playing .workshops-icon-col { transform: translateY(7.5px); } WorkshopsLive sessions .evals-icon-row { transition: transform 350ms ease-in-out; } .evals-icon-playing .evals-icon-row-down { transform: translateY(5px); } .evals-icon-playing .evals-icon-row-up { transform: translateY(-5px); } Eval libraryOpen source researchFoundationsLearn to evalEncyclopediaEval conceptsProductObserveEvaluateDiscoverResourcesDocsBlogWorkshopsEval libraryFoundationsEncyclopediaCustomersPricingContact usSign inSign upEvaluate your AI development loop.Not all agents are created with the same quality and feedback loops. Find out where your team ranks given today's leading agent development workflows.Start assessment01Describe the productName the agent, workflow, or AI feature you want to improve.02Map the loopMark the quality practices your team already has in motion.03Set confidenceShare how confidently you can catch regressions today.React

Discovered: Last checked: Content changed:
ArticleInternal link

.evals-icon-row { transition: transform 350ms ease-in-out; } .evals-icon-playing .evals-icon-row-down { transform: translateY(5px); } .evals-icon-playing .evals-icon-row-up { transform: translateY(-5px); } Eval libraryOpen source research

https://www.braintrust.dev/evals

Open original page

Independent eval research on models, agents, and cost. The work is open source, so you can check the methodology or rerun a study on your own data.

Language: en
Indexed excerpt

ProductResourcesCustomersPricingContact usSign inSign upObserveTrace everythingEvaluateTest what shipsDiscoverFind behaviorsDocsStart buildingBlogInsights and updates .workshops-icon-col { transition: transform 400ms cubic-bezier(0.65, 0, 0.35, 1); } .workshops-icon-playing .workshops-icon-col { transform: translateY(7.5px); } WorkshopsLive sessions .evals-icon-row { transition: transform 350ms ease-in-out; } .evals-icon-playing .evals-icon-row-down { transform: translateY(5px); } .evals-icon-playing .evals-icon-row-up { transform: translateY(-5px); } Eval libraryOpen source researchFoundationsLearn to evalEncyclopediaEval conceptsProductObserveEvaluateDiscoverResourcesDocsBlogWorkshopsEval libraryFoundationsEncyclopediaCustomersPricingContact usSign inSign upIndependent eval research. Open methodology, published datasets, and the tools to run your own studies.ResearchVideosSkillsNewsletterOriginal, open-source studiesModel comparisons, agent setups, cost, and modalities, measured on real tasks. Read the methodology and results, or rerun a study on your own data.9 September 2026Moonshot vs Fireworks for Kimi K3 frontend agentsKimi K3 ran through Moonshot and Fireworks in 30 direct

Discovered: Last checked: Content changed:
ArticleInternal link

EncyclopediaEval concepts

https://www.braintrust.dev/encyclopedia

Open original page

A complete alphabetical index of Encyclopedia Evalica terms and definitions.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleInternal link

.workshops-icon-col { transition: transform 400ms cubic-bezier(0.65, 0, 0.35, 1); } .workshops-icon-playing .workshops-icon-col { transform: translateY(7.5px); } WorkshopsLive sessions

https://www.braintrust.dev/workshops

Open original page

Learn the foundations of evals, discover how to measure AI quality, and build a repeatable eval workflow.

Language: en
Indexed excerpt

ProductResourcesCustomersPricingContact usSign inSign upObserveTrace everythingEvaluateTest what shipsDiscoverFind behaviorsDocsStart buildingBlogInsights and updates .workshops-icon-col { transition: transform 400ms cubic-bezier(0.65, 0, 0.35, 1); } .workshops-icon-playing .workshops-icon-col { transform: translateY(7.5px); } WorkshopsLive sessions .evals-icon-row { transition: transform 350ms ease-in-out; } .evals-icon-playing .evals-icon-row-down { transform: translateY(5px); } .evals-icon-playing .evals-icon-row-up { transform: translateY(-5px); } Eval libraryOpen source researchFoundationsLearn to evalEncyclopediaEval conceptsProductObserveEvaluateDiscoverResourcesDocsBlogWorkshopsEval libraryFoundationsEncyclopediaCustomersPricingContact usSign inSign upRegister for live workshops with the Braintrust team, or rewatch past sessions on demand.All audiencesPMsEngineersUpcomingEvals foundations7 October 202610:00 AM PTDesigned for new Braintrust users, this workshop will cover the core foundations of evals and how to get started with Braintrust.PMsKeep up with new workshopsGet notified when new events, live workshops, and on-demand sessions are added.Email addressNotify meOn Dema

Discovered: Last checked: Content changed:
ArticleInternal link

In-depth articles and insights

https://www.braintrust.dev/articles

Open original page

In-depth articles and insights about AI evaluation, development best practices, and technical deep dives from the Braintrust team.

Language: en
Indexed excerpt

ProductResourcesCustomersPricingContact usSign inSign upObserveTrace everythingEvaluateTest what shipsDiscoverFind behaviorsDocsStart buildingBlogInsights and updates .workshops-icon-col { transition: transform 400ms cubic-bezier(0.65, 0, 0.35, 1); } .workshops-icon-playing .workshops-icon-col { transform: translateY(7.5px); } WorkshopsLive sessions .evals-icon-row { transition: transform 350ms ease-in-out; } .evals-icon-playing .evals-icon-row-down { transform: translateY(5px); } .evals-icon-playing .evals-icon-row-up { transform: translateY(-5px); } Eval libraryOpen source researchFoundationsLearn to evalEncyclopediaEval conceptsProductObserveEvaluateDiscoverResourcesDocsBlogWorkshopsEval libraryFoundationsEncyclopediaCustomersPricingContact usSign inSign upGuides for shipping quality AIBraintrust works with customers building AI products at scale.These guides distill the patterns that work. Instrument observability to understand how AI behaves in production. Use evals to improve your AI productsStart buildingLatest articlesRSSBest no-code AI agent builders in 2026Compare the best no-code AI agent builders in 2026. See how Sim AI, Lindy, Relevance AI, Stack AI, Gumloop, and Barde

Discovered: Last checked: Content changed: