文章 Sitemap 2026/10/7
2026 10 07 evals foundations https://www.braintrust.dev/workshops/2026-10-07-evals-foundations
前往原始頁面 ↗ Designed for new Braintrust users, this workshop will cover the core foundations of evals and how to get started with Braintrust.
語言: en
索引內容摘要 WorkshopsEvals foundations7 October 202610:00 AM PTPMsJess WangDeveloper AdvocateRegister for the workshopDesigned for new Braintrust users, this workshop will cover the core foundations of evals and how to get started with Braintrust. Jess will walk through the practical pieces of the eval workflow: moving from a question about your AI system to experiments, results, and feedback you can use to improve it. Bring a system you are building, a quality question you want to answer, or simply your curiosity. You will leave with a clearer picture of how to measure and improve agent quality, plus a practical next step to take. What you’ll learn How Braintrust groups production traces into named failure patterns you can act on How to filter a failure cluster into a labeled dataset How to write an eval that targets a specific failure pattern and validate the fix How to run a repeatable diagnose-to-eval workflow in Braintrust Recent on-demand workshopsEvals foundations9 September 202610:00 AM PTDesigned for new Braintrust users, this workshop will cover the core foundations of evals and how to get started with Braintrust.PMsHow financial services teams ship AI agents to production26 August 2…
首次發現: 2026/9/10 13:30:50 最近檢查: 2026/9/20 11:00:52 內容更新: 2026/9/19 18:30:17
文章 Sitemap 2026/9/20
Golden dataset https://www.braintrust.dev/encyclopedia/golden-dataset
前往原始頁面 ↗ A curated, high-quality evaldataset that serves as the canonical benchmark for a team's use case. It represents critical functionality and known failure modes.
語言: en
索引內容摘要 The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj…
首次發現: 2026/9/10 13:30:48 最近檢查: 2026/9/20 12:31:10 內容更新: 2026/9/20 12:31:10
文章 Sitemap 2026/9/20
Groundedness https://www.braintrust.dev/encyclopedia/groundedness
前往原始頁面 ↗ A scoring dimension that measures whether a generated answer is grounded in and directly supported by evidence in the provided context. It's especially important for RAG and policy-heavy domains.
語言: en
索引內容摘要 The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj…
首次發現: 2026/9/10 13:30:48 最近檢查: 2026/9/20 12:31:10 內容更新: 2026/9/20 12:31:10
文章 Sitemap 2026/9/20
Feedback signal https://www.braintrust.dev/encyclopedia/feedback-signal
前往原始頁面 ↗ Any quantitative or categorical signal used for iteration (thumbs up/down, labels, scores, outcomes). Strong feedback signals help prioritize what to fix next.
語言: en
索引內容摘要 The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj…
首次發現: 2026/9/10 13:30:48 最近檢查: 2026/9/20 12:31:10 內容更新: 2026/9/20 12:31:10
文章 Sitemap 2026/9/20
Function (deployed) https://www.braintrust.dev/encyclopedia/function-deployed
前往原始頁面 ↗ A versioned task/function deployed behind an endpoint (often the unit you evaluate and compare). Deployed functions let teams connect eval outcomes to real production rollouts.
語言: en
索引內容摘要 The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj…
首次發現: 2026/9/10 13:30:48 最近檢查: 2026/9/20 12:31:10 內容更新: 2026/9/20 12:31:10
文章 Sitemap 2026/9/20
Fallback https://www.braintrust.dev/encyclopedia/fallback
前往原始頁面 ↗ The process of retrying a failed LLM request with an alternative provider or model. Fallbacks improve reliability when providers are rate-limited or unstable.
語言: en
索引內容摘要 The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj…
首次發現: 2026/9/10 13:30:48 最近檢查: 2026/9/20 12:31:10 內容更新: 2026/9/20 12:31:10
文章 Sitemap 2026/9/20
Failure mode https://www.braintrust.dev/encyclopedia/failure-mode
前往原始頁面 ↗ A specific, recurring pattern of incorrect or undesired output. Identifying failure modes is the starting point for targeted improvement.
語言: en
索引內容摘要 The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj…
首次發現: 2026/9/10 13:30:48 最近檢查: 2026/9/20 12:31:10 內容更新: 2026/9/20 12:31:10
文章 Sitemap 2026/9/20
Feedback loop https://www.braintrust.dev/encyclopedia/feedback-loop
前往原始頁面 ↗ The continuous cycle connecting production traces to eval datasets to system improvements and back to production.
語言: en
索引內容摘要 The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj…
首次發現: 2026/9/10 13:30:48 最近檢查: 2026/9/20 12:31:10 內容更新: 2026/9/20 12:31:10
文章 Sitemap 2026/9/20
flush() https://www.braintrust.dev/encyclopedia/flush
前往原始頁面 ↗ An SDK method that forces all buffered dataset records or trace spans to be written to Braintrust immediately. This is useful in short-lived jobs (like serverless) where buffered data might not be sent before the process exits.
語言: en
索引內容摘要 The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj…
首次發現: 2026/9/10 13:30:48 最近檢查: 2026/9/20 12:31:10 內容更新: 2026/9/20 12:31:10
文章 Sitemap 2026/9/20
Experiment https://www.braintrust.dev/encyclopedia/experiment
前往原始頁面 ↗ An immutable snapshot of a single eval run, including the dataset, task configuration, scores, and outputs. Experiments make results reproducible and shareable.
語言: en
索引內容摘要 The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj…
首次發現: 2026/9/10 13:30:48 最近檢查: 2026/9/20 12:31:10 內容更新: 2026/9/20 12:31:10
文章 Sitemap 2026/9/20
Eval harness https://www.braintrust.dev/encyclopedia/eval-harness
前往原始頁面 ↗ The code and infrastructure that runs an eval end-to-end (dataset loading, task execution, scoring, reporting). A harness makes evals repeatable across environments.
語言: en
索引內容摘要 The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj…
首次發現: 2026/9/10 13:30:48 最近檢查: 2026/9/20 12:31:10 內容更新: 2026/9/20 12:31:10
文章 Sitemap 2026/9/20
Factuality https://www.braintrust.dev/encyclopedia/factuality
前往原始頁面 ↗ A quality dimension measuring whether statements are factually correct. Factuality can be high even when an answer is poorly structured or irrelevant.
語言: en
索引內容摘要 The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj…
首次發現: 2026/9/10 13:30:48 最近檢查: 2026/9/20 12:31:10 內容更新: 2026/9/20 12:31:10
文章 Sitemap 2026/9/20
Faithfulness https://www.braintrust.dev/encyclopedia/faithfulness
前往原始頁面 ↗ A quality dimension measuring whether an output is supported by the provided context (i.e., no contradictions or unsupported claims). Faithfulness is often used to measure hallucinations against context.
語言: en
索引內容摘要 The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj…
首次發現: 2026/9/10 13:30:48 最近檢查: 2026/9/20 12:31:10 內容更新: 2026/9/20 12:31:10
文章 Sitemap 2026/9/20
Expected output (ground truth) https://www.braintrust.dev/encyclopedia/expected-output-ground-truth
前往原始頁面 ↗ The reference answer/behavior a dataset record defines as correct (when doing reference-based evals). Ground truth can be human-written, programmatically derived, or imported from production.
語言: en
索引內容摘要 The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj…
首次發現: 2026/9/10 13:30:48 最近檢查: 2026/9/20 12:31:10 內容更新: 2026/9/20 12:31:10
文章 Sitemap 2026/9/20
Drift https://www.braintrust.dev/encyclopedia/drift
前往原始頁面 ↗ A gradual shift in the distribution of inputs, outputs, or scores over time. Drift can signal changes in user behavior, model degradation, or upstream product changes.
語言: en
索引內容摘要 The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj…
首次發現: 2026/9/10 13:30:48 最近檢查: 2026/9/20 12:31:10 內容更新: 2026/9/20 12:31:10
文章 Sitemap 2026/9/20
Error rate https://www.braintrust.dev/encyclopedia/error-rate
前往原始頁面 ↗ The fraction of requests that fail due to system/tool/provider errors (distinct from semantic failures). Tracking error rate helps separate reliability problems from quality problems.
語言: en
索引內容摘要 The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj…
首次發現: 2026/9/10 13:30:48 最近檢查: 2026/9/20 12:31:10 內容更新: 2026/9/20 12:31:10
文章 Sitemap 2026/9/20
Eval leakage https://www.braintrust.dev/encyclopedia/eval-leakage
前往原始頁面 ↗ When eval data unintentionally influences the system being evaluated (e.g., test cases appear in training or prompts). Leakage can inflate scores without improving real-world performance.
語言: en
索引內容摘要 The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj…
首次發現: 2026/9/10 13:30:48 最近檢查: 2026/9/20 12:31:10 內容更新: 2026/9/20 12:31:10
文章 Sitemap 2026/9/20
Data flywheel https://www.braintrust.dev/encyclopedia/data-flywheel
前往原始頁面 ↗ A feedback loop in which production data improves eval datasets, which improves the system, which in turn produces better production data. A strong flywheel turns "bugs" into durable eval coverage.
語言: en
索引內容摘要 The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj…
首次發現: 2026/9/10 13:30:48 最近檢查: 2026/9/20 12:31:10 內容更新: 2026/9/20 12:31:10
文章 Sitemap 2026/9/20
Deep search https://www.braintrust.dev/encyclopedia/deep-search
前往原始頁面 ↗ Searching across large volumes of traces/datasets with filters and semantic search to find patterns and outliers. It's especially useful for finding "nearby" failures that share structure but not keywords.
語言: en
索引內容摘要 The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj…
首次發現: 2026/9/10 13:30:48 最近檢查: 2026/9/20 12:31:10 內容更新: 2026/9/20 12:31:10
文章 Sitemap 2026/9/20
Dataset record https://www.braintrust.dev/encyclopedia/dataset-record
前往原始頁面 ↗ A single row in a dataset, consisting of an input, an expected output, and optional metadata. Records are the unit you score and analyze.
語言: en
索引內容摘要 The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj…
首次發現: 2026/9/10 13:30:48 最近檢查: 2026/9/20 12:31:10 內容更新: 2026/9/20 12:31:10
文章 Sitemap 2026/9/20
Edge case https://www.braintrust.dev/encyclopedia/edge-case
前往原始頁面 ↗ An input that lies at the boundary of expected behavior. Edge cases are rare but disproportionately important to handle correctly.
語言: en
索引內容摘要 The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj…
首次發現: 2026/9/10 13:30:48 最近檢查: 2026/9/20 12:31:09 內容更新: 2026/9/20 12:31:09
文章 Sitemap 2026/9/20
Dashboard https://www.braintrust.dev/encyclopedia/dashboard
前往原始頁面 ↗ A configurable view that tracks key metrics over time using charts and aggregations, covering latency, cost, error rates, and quality scores. Dashboards make it easy to spot trends and regressions without querying raw traces.
語言: en
索引內容摘要 The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj…
首次發現: 2026/9/10 13:30:48 最近檢查: 2026/9/20 12:31:09 內容更新: 2026/9/20 12:31:09
文章 站內連結 2026/9/20
The principles of Braintrust https://www.braintrust.dev/manifesto
前往原始頁面 ↗ Braintrust's principles, beliefs, and practical guidance for building high-quality evals.
語言: en
索引內容摘要 Evals and observability are AI infrastructureThe eval is as important as the model, harness, etc. It should be treated as part of the stack, not an afterthought.If you’re building agents, you need to consider the infrastructure stack the same way you would for any software. What cloud provider, database, what IaC framework? What experiments, which scorers, which rubrics?Agents are now in production, supporting products used by millions. We are past the point of toy demos and neat experiments. Serious products need serious infrastructure. And real infrastructure is observable.
首次發現: 2026/9/10 13:30:48 最近檢查: 2026/9/20 12:30:59 內容更新: 2026/9/20 10:00:50
文章 站內連結 2026/9/20
Pricing https://www.braintrust.dev/pricing
前往原始頁面 ↗ Start building AI products for free with Braintrust. Transparent pricing for AI evaluation, monitoring, and observability. No hidden fees, pay as you scale.
語言: en
索引內容摘要 ProductResourcesCustomersPricingContact usSign inSign upObserveTrace everythingEvaluateTest what shipsDiscoverFind behaviorsDocsStart buildingBlogInsights and updates .workshops-icon-col { transition: transform 400ms cubic-bezier(0.65, 0, 0.35, 1); } .workshops-icon-playing .workshops-icon-col { transform: translateY(7.5px); } WorkshopsLive sessions .evals-icon-row { transition: transform 350ms ease-in-out; } .evals-icon-playing .evals-icon-row-down { transform: translateY(5px); } .evals-icon-playing .evals-icon-row-up { transform: translateY(-5px); } Eval libraryOpen source researchFoundationsLearn to evalEncyclopediaEval conceptsProductObserveEvaluateDiscoverResourcesDocsBlogWorkshopsEval libraryFoundationsEncyclopediaCustomersPricingContact usSign inSign upPredictable pricing. Designed to scale.Start now. Instrument when you’re ready.No credit card required. Free usage is included for traces, evals, and your whole team.Sign up Start building for freeModel creditsProcessed dataScoresData retentionFeaturesStarterFor everyone$0 / month$10 credits+ tok rates1 GB processed data+ $4/GB10k scores+ $2.50/1k14-day retentionUnlimited users, projects, datasets, playgrounds, and experiments…
首次發現: 2026/9/10 13:30:48 最近檢查: 2026/9/20 12:30:59 內容更新: 2026/9/20 10:00:44
文章 站內連結 2026/9/20
Take the eval maturity assessment https://www.braintrust.dev/assessment
前往原始頁面 ↗ Find gaps in your AI development loop and get an AI quality roadmap.
語言: en
索引內容摘要 ProductResourcesCustomersPricingContact usSign inSign upObserveTrace everythingEvaluateTest what shipsDiscoverFind behaviorsDocsStart buildingBlogInsights and updates .workshops-icon-col { transition: transform 400ms cubic-bezier(0.65, 0, 0.35, 1); } .workshops-icon-playing .workshops-icon-col { transform: translateY(7.5px); } WorkshopsLive sessions .evals-icon-row { transition: transform 350ms ease-in-out; } .evals-icon-playing .evals-icon-row-down { transform: translateY(5px); } .evals-icon-playing .evals-icon-row-up { transform: translateY(-5px); } Eval libraryOpen source researchFoundationsLearn to evalEncyclopediaEval conceptsProductObserveEvaluateDiscoverResourcesDocsBlogWorkshopsEval libraryFoundationsEncyclopediaCustomersPricingContact usSign inSign upEvaluate your AI development loop.Not all agents are created with the same quality and feedback loops. Find out where your team ranks given today's leading agent development workflows.Start assessment01Describe the productName the agent, workflow, or AI feature you want to improve.02Map the loopMark the quality practices your team already has in motion.03Set confidenceShare how confidently you can catch regressions today.React…
首次發現: 2026/9/10 13:30:48 最近檢查: 2026/9/20 12:30:59 內容更新: 2026/9/20 10:00:44