SEARWEB SITE INDEX

Braintrust · Indexed content

braintrust.dev

Explore internal pages, articles and content excerpts discovered from this site’s public sources.

Internal links
462
Articles
393
Last indexed
20‏/9‏/2026، 12:31:10 م
462 indexed items
ArticleSitemap

2026 10 07 evals foundations

https://www.braintrust.dev/workshops/2026-10-07-evals-foundations

Open original page

Designed for new Braintrust users, this workshop will cover the core foundations of evals and how to get started with Braintrust.

Language: en
Indexed excerpt

WorkshopsEvals foundations7 October 202610:00 AM PTPMsJess WangDeveloper AdvocateRegister for the workshopDesigned for new Braintrust users, this workshop will cover the core foundations of evals and how to get started with Braintrust. Jess will walk through the practical pieces of the eval workflow: moving from a question about your AI system to experiments, results, and feedback you can use to improve it. Bring a system you are building, a quality question you want to answer, or simply your curiosity. You will leave with a clearer picture of how to measure and improve agent quality, plus a practical next step to take. What you’ll learn How Braintrust groups production traces into named failure patterns you can act on How to filter a failure cluster into a labeled dataset How to write an eval that targets a specific failure pattern and validate the fix How to run a repeatable diagnose-to-eval workflow in Braintrust Recent on-demand workshopsEvals foundations9 September 202610:00 AM PTDesigned for new Braintrust users, this workshop will cover the core foundations of evals and how to get started with Braintrust.PMsHow financial services teams ship AI agents to production26 August 2

Discovered: Last checked: Content changed:
ArticleSitemap

Golden dataset

https://www.braintrust.dev/encyclopedia/golden-dataset

Open original page

A curated, high-quality evaldataset that serves as the canonical benchmark for a team's use case. It represents critical functionality and known failure modes.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Groundedness

https://www.braintrust.dev/encyclopedia/groundedness

Open original page

A scoring dimension that measures whether a generated answer is grounded in and directly supported by evidence in the provided context. It's especially important for RAG and policy-heavy domains.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Feedback signal

https://www.braintrust.dev/encyclopedia/feedback-signal

Open original page

Any quantitative or categorical signal used for iteration (thumbs up/down, labels, scores, outcomes). Strong feedback signals help prioritize what to fix next.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Function (deployed)

https://www.braintrust.dev/encyclopedia/function-deployed

Open original page

A versioned task/function deployed behind an endpoint (often the unit you evaluate and compare). Deployed functions let teams connect eval outcomes to real production rollouts.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Fallback

https://www.braintrust.dev/encyclopedia/fallback

Open original page

The process of retrying a failed LLM request with an alternative provider or model. Fallbacks improve reliability when providers are rate-limited or unstable.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Failure mode

https://www.braintrust.dev/encyclopedia/failure-mode

Open original page

A specific, recurring pattern of incorrect or undesired output. Identifying failure modes is the starting point for targeted improvement.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Feedback loop

https://www.braintrust.dev/encyclopedia/feedback-loop

Open original page

The continuous cycle connecting production traces to eval datasets to system improvements and back to production.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

flush()

https://www.braintrust.dev/encyclopedia/flush

Open original page

An SDK method that forces all buffered dataset records or trace spans to be written to Braintrust immediately. This is useful in short-lived jobs (like serverless) where buffered data might not be sent before the process exits.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Experiment

https://www.braintrust.dev/encyclopedia/experiment

Open original page

An immutable snapshot of a single eval run, including the dataset, task configuration, scores, and outputs. Experiments make results reproducible and shareable.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Eval harness

https://www.braintrust.dev/encyclopedia/eval-harness

Open original page

The code and infrastructure that runs an eval end-to-end (dataset loading, task execution, scoring, reporting). A harness makes evals repeatable across environments.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Factuality

https://www.braintrust.dev/encyclopedia/factuality

Open original page

A quality dimension measuring whether statements are factually correct. Factuality can be high even when an answer is poorly structured or irrelevant.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Faithfulness

https://www.braintrust.dev/encyclopedia/faithfulness

Open original page

A quality dimension measuring whether an output is supported by the provided context (i.e., no contradictions or unsupported claims). Faithfulness is often used to measure hallucinations against context.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Expected output (ground truth)

https://www.braintrust.dev/encyclopedia/expected-output-ground-truth

Open original page

The reference answer/behavior a dataset record defines as correct (when doing reference-based evals). Ground truth can be human-written, programmatically derived, or imported from production.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Drift

https://www.braintrust.dev/encyclopedia/drift

Open original page

A gradual shift in the distribution of inputs, outputs, or scores over time. Drift can signal changes in user behavior, model degradation, or upstream product changes.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Error rate

https://www.braintrust.dev/encyclopedia/error-rate

Open original page

The fraction of requests that fail due to system/tool/provider errors (distinct from semantic failures). Tracking error rate helps separate reliability problems from quality problems.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Eval leakage

https://www.braintrust.dev/encyclopedia/eval-leakage

Open original page

When eval data unintentionally influences the system being evaluated (e.g., test cases appear in training or prompts). Leakage can inflate scores without improving real-world performance.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Data flywheel

https://www.braintrust.dev/encyclopedia/data-flywheel

Open original page

A feedback loop in which production data improves eval datasets, which improves the system, which in turn produces better production data. A strong flywheel turns "bugs" into durable eval coverage.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Deep search

https://www.braintrust.dev/encyclopedia/deep-search

Open original page

Searching across large volumes of traces/datasets with filters and semantic search to find patterns and outliers. It's especially useful for finding "nearby" failures that share structure but not keywords.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Dataset record

https://www.braintrust.dev/encyclopedia/dataset-record

Open original page

A single row in a dataset, consisting of an input, an expected output, and optional metadata. Records are the unit you score and analyze.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Edge case

https://www.braintrust.dev/encyclopedia/edge-case

Open original page

An input that lies at the boundary of expected behavior. Edge cases are rare but disproportionately important to handle correctly.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleSitemap

Dashboard

https://www.braintrust.dev/encyclopedia/dashboard

Open original page

A configurable view that tracks key metrics over time using charts and aggregations, covering latency, cost, error rates, and quality scores. Dashboards make it easy to spot trends and regressions without querying raw traces.

Language: en
Indexed excerpt

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

Discovered: Last checked: Content changed:
ArticleInternal link

The principles of Braintrust

https://www.braintrust.dev/manifesto

Open original page

Braintrust's principles, beliefs, and practical guidance for building high-quality evals.

Language: en
Indexed excerpt

Evals and observability are AI infrastructureThe eval is as important as the model, harness, etc. It should be treated as part of the stack, not an afterthought.If you’re building agents, you need to consider the infrastructure stack the same way you would for any software. What cloud provider, database, what IaC framework? What experiments, which scorers, which rubrics?Agents are now in production, supporting products used by millions. We are past the point of toy demos and neat experiments. Serious products need serious infrastructure. And real infrastructure is observable.

Discovered: Last checked: Content changed:
ArticleInternal link

Pricing

https://www.braintrust.dev/pricing

Open original page

Start building AI products for free with Braintrust. Transparent pricing for AI evaluation, monitoring, and observability. No hidden fees, pay as you scale.

Language: en
Indexed excerpt

ProductResourcesCustomersPricingContact usSign inSign upObserveTrace everythingEvaluateTest what shipsDiscoverFind behaviorsDocsStart buildingBlogInsights and updates .workshops-icon-col { transition: transform 400ms cubic-bezier(0.65, 0, 0.35, 1); } .workshops-icon-playing .workshops-icon-col { transform: translateY(7.5px); } WorkshopsLive sessions .evals-icon-row { transition: transform 350ms ease-in-out; } .evals-icon-playing .evals-icon-row-down { transform: translateY(5px); } .evals-icon-playing .evals-icon-row-up { transform: translateY(-5px); } Eval libraryOpen source researchFoundationsLearn to evalEncyclopediaEval conceptsProductObserveEvaluateDiscoverResourcesDocsBlogWorkshopsEval libraryFoundationsEncyclopediaCustomersPricingContact usSign inSign upPredictable pricing. Designed to scale.Start now. Instrument when you’re ready.No credit card required. Free usage is included for traces, evals, and your whole team.Sign up Start building for freeModel creditsProcessed dataScoresData retentionFeaturesStarterFor everyone$0 / month$10 credits+ tok rates1 GB processed data+ $4/GB10k scores+ $2.50/1k14-day retentionUnlimited users, projects, datasets, playgrounds, and experiments

Discovered: Last checked: Content changed:
ArticleInternal link

Take the eval maturity assessment

https://www.braintrust.dev/assessment

Open original page

Find gaps in your AI development loop and get an AI quality roadmap.

Language: en
Indexed excerpt

ProductResourcesCustomersPricingContact usSign inSign upObserveTrace everythingEvaluateTest what shipsDiscoverFind behaviorsDocsStart buildingBlogInsights and updates .workshops-icon-col { transition: transform 400ms cubic-bezier(0.65, 0, 0.35, 1); } .workshops-icon-playing .workshops-icon-col { transform: translateY(7.5px); } WorkshopsLive sessions .evals-icon-row { transition: transform 350ms ease-in-out; } .evals-icon-playing .evals-icon-row-down { transform: translateY(5px); } .evals-icon-playing .evals-icon-row-up { transform: translateY(-5px); } Eval libraryOpen source researchFoundationsLearn to evalEncyclopediaEval conceptsProductObserveEvaluateDiscoverResourcesDocsBlogWorkshopsEval libraryFoundationsEncyclopediaCustomersPricingContact usSign inSign upEvaluate your AI development loop.Not all agents are created with the same quality and feedback loops. Find out where your team ranks given today's leading agent development workflows.Start assessment01Describe the productName the agent, workflow, or AI feature you want to improve.02Map the loopMark the quality practices your team already has in motion.03Set confidenceShare how confidently you can catch regressions today.React

Discovered: Last checked: Content changed: