SEARWEB 網站索引

Braintrust · 索引內容

braintrust.dev

查看 SEARWEB 從此網站公開來源發現的站內頁面、文章和內容摘要。

站內連結
462
文章
393
最近索引
2026/9/20 下午12:31:10
462 項索引內容
文章Sitemap

2026 10 07 evals foundations

https://www.braintrust.dev/workshops/2026-10-07-evals-foundations

前往原始頁面

Designed for new Braintrust users, this workshop will cover the core foundations of evals and how to get started with Braintrust.

語言: en
索引內容摘要

WorkshopsEvals foundations7 October 202610:00 AM PTPMsJess WangDeveloper AdvocateRegister for the workshopDesigned for new Braintrust users, this workshop will cover the core foundations of evals and how to get started with Braintrust. Jess will walk through the practical pieces of the eval workflow: moving from a question about your AI system to experiments, results, and feedback you can use to improve it. Bring a system you are building, a quality question you want to answer, or simply your curiosity. You will leave with a clearer picture of how to measure and improve agent quality, plus a practical next step to take. What you’ll learn How Braintrust groups production traces into named failure patterns you can act on How to filter a failure cluster into a labeled dataset How to write an eval that targets a specific failure pattern and validate the fix How to run a repeatable diagnose-to-eval workflow in Braintrust Recent on-demand workshopsEvals foundations9 September 202610:00 AM PTDesigned for new Braintrust users, this workshop will cover the core foundations of evals and how to get started with Braintrust.PMsHow financial services teams ship AI agents to production26 August 2

首次發現: 最近檢查: 內容更新:
文章Sitemap

Golden dataset

https://www.braintrust.dev/encyclopedia/golden-dataset

前往原始頁面

A curated, high-quality evaldataset that serves as the canonical benchmark for a team's use case. It represents critical functionality and known failure modes.

語言: en
索引內容摘要

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

首次發現: 最近檢查: 內容更新:
文章Sitemap

Groundedness

https://www.braintrust.dev/encyclopedia/groundedness

前往原始頁面

A scoring dimension that measures whether a generated answer is grounded in and directly supported by evidence in the provided context. It's especially important for RAG and policy-heavy domains.

語言: en
索引內容摘要

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

首次發現: 最近檢查: 內容更新:
文章Sitemap

Feedback signal

https://www.braintrust.dev/encyclopedia/feedback-signal

前往原始頁面

Any quantitative or categorical signal used for iteration (thumbs up/down, labels, scores, outcomes). Strong feedback signals help prioritize what to fix next.

語言: en
索引內容摘要

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

首次發現: 最近檢查: 內容更新:
文章Sitemap

Function (deployed)

https://www.braintrust.dev/encyclopedia/function-deployed

前往原始頁面

A versioned task/function deployed behind an endpoint (often the unit you evaluate and compare). Deployed functions let teams connect eval outcomes to real production rollouts.

語言: en
索引內容摘要

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

首次發現: 最近檢查: 內容更新:
文章Sitemap

Fallback

https://www.braintrust.dev/encyclopedia/fallback

前往原始頁面

The process of retrying a failed LLM request with an alternative provider or model. Fallbacks improve reliability when providers are rate-limited or unstable.

語言: en
索引內容摘要

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

首次發現: 最近檢查: 內容更新:
文章Sitemap

Failure mode

https://www.braintrust.dev/encyclopedia/failure-mode

前往原始頁面

A specific, recurring pattern of incorrect or undesired output. Identifying failure modes is the starting point for targeted improvement.

語言: en
索引內容摘要

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

首次發現: 最近檢查: 內容更新:
文章Sitemap

Feedback loop

https://www.braintrust.dev/encyclopedia/feedback-loop

前往原始頁面

The continuous cycle connecting production traces to eval datasets to system improvements and back to production.

語言: en
索引內容摘要

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

首次發現: 最近檢查: 內容更新:
文章Sitemap

flush()

https://www.braintrust.dev/encyclopedia/flush

前往原始頁面

An SDK method that forces all buffered dataset records or trace spans to be written to Braintrust immediately. This is useful in short-lived jobs (like serverless) where buffered data might not be sent before the process exits.

語言: en
索引內容摘要

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

首次發現: 最近檢查: 內容更新:
文章Sitemap

Experiment

https://www.braintrust.dev/encyclopedia/experiment

前往原始頁面

An immutable snapshot of a single eval run, including the dataset, task configuration, scores, and outputs. Experiments make results reproducible and shareable.

語言: en
索引內容摘要

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

首次發現: 最近檢查: 內容更新:
文章Sitemap

Eval harness

https://www.braintrust.dev/encyclopedia/eval-harness

前往原始頁面

The code and infrastructure that runs an eval end-to-end (dataset loading, task execution, scoring, reporting). A harness makes evals repeatable across environments.

語言: en
索引內容摘要

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

首次發現: 最近檢查: 內容更新:
文章Sitemap

Factuality

https://www.braintrust.dev/encyclopedia/factuality

前往原始頁面

A quality dimension measuring whether statements are factually correct. Factuality can be high even when an answer is poorly structured or irrelevant.

語言: en
索引內容摘要

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

首次發現: 最近檢查: 內容更新:
文章Sitemap

Faithfulness

https://www.braintrust.dev/encyclopedia/faithfulness

前往原始頁面

A quality dimension measuring whether an output is supported by the provided context (i.e., no contradictions or unsupported claims). Faithfulness is often used to measure hallucinations against context.

語言: en
索引內容摘要

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

首次發現: 最近檢查: 內容更新:
文章Sitemap

Expected output (ground truth)

https://www.braintrust.dev/encyclopedia/expected-output-ground-truth

前往原始頁面

The reference answer/behavior a dataset record defines as correct (when doing reference-based evals). Ground truth can be human-written, programmatically derived, or imported from production.

語言: en
索引內容摘要

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

首次發現: 最近檢查: 內容更新:
文章Sitemap

Drift

https://www.braintrust.dev/encyclopedia/drift

前往原始頁面

A gradual shift in the distribution of inputs, outputs, or scores over time. Drift can signal changes in user behavior, model degradation, or upstream product changes.

語言: en
索引內容摘要

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

首次發現: 最近檢查: 內容更新:
文章Sitemap

Error rate

https://www.braintrust.dev/encyclopedia/error-rate

前往原始頁面

The fraction of requests that fail due to system/tool/provider errors (distinct from semantic failures). Tracking error rate helps separate reliability problems from quality problems.

語言: en
索引內容摘要

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

首次發現: 最近檢查: 內容更新:
文章Sitemap

Eval leakage

https://www.braintrust.dev/encyclopedia/eval-leakage

前往原始頁面

When eval data unintentionally influences the system being evaluated (e.g., test cases appear in training or prompts). Leakage can inflate scores without improving real-world performance.

語言: en
索引內容摘要

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

首次發現: 最近檢查: 內容更新:
文章Sitemap

Data flywheel

https://www.braintrust.dev/encyclopedia/data-flywheel

前往原始頁面

A feedback loop in which production data improves eval datasets, which improves the system, which in turn produces better production data. A strong flywheel turns "bugs" into durable eval coverage.

語言: en
索引內容摘要

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

首次發現: 最近檢查: 內容更新:
文章Sitemap

Deep search

https://www.braintrust.dev/encyclopedia/deep-search

前往原始頁面

Searching across large volumes of traces/datasets with filters and semantic search to find patterns and outliers. It's especially useful for finding "nearby" failures that share structure but not keywords.

語言: en
索引內容摘要

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

首次發現: 最近檢查: 內容更新:
文章Sitemap

Dataset record

https://www.braintrust.dev/encyclopedia/dataset-record

前往原始頁面

A single row in a dataset, consisting of an input, an expected output, and optional metadata. Records are the unit you score and analyze.

語言: en
索引內容摘要

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

首次發現: 最近檢查: 內容更新:
文章Sitemap

Edge case

https://www.braintrust.dev/encyclopedia/edge-case

前往原始頁面

An input that lies at the boundary of expected behavior. Edge cases are rare but disproportionately important to handle correctly.

語言: en
索引內容摘要

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

首次發現: 最近檢查: 內容更新:
文章Sitemap

Dashboard

https://www.braintrust.dev/encyclopedia/dashboard

前往原始頁面

A configurable view that tracks key metrics over time using charts and aggregations, covering latency, cost, error rates, and quality scores. Dashboards make it easy to spot trends and regressions without querying raw traces.

語言: en
索引內容摘要

The Eval ManifestoEncyclopedia EvalicaAAbsolute scoringActive observabilityAdversarial examplesAgentAgent observabilityAlert / thresholdAlignmentAnnotationAnnotation schemaAttachmentBBaselineBaseline experimentBenchmarkBrainstoreCCachingCalibrationCI/CD integrationCoherenceConfidence intervalCoverageDDashboardData flywheelDatasetDataset recordDeep searchDriftEEvalEdge caseError rateEval harnessEval leakageExpected output (ground truth)ExperimentFFactualityFailure modeFaithfulnessFallbackFeedback loopFeedback signalflush()Function (deployed)GGatewayGolden datasetGroundednessHHallucinationHuman reviewHybrid deploymentIInputInstrumentationInter-annotator agreement (IAA)LLabelLLM-as-a-judgeLogsLoopMMetadataModel comparisonModel driftMultimodalMultimodal datasetNNon-determinismOOffline evaluationOnline evaluation (production scoring)PP50 / P95 / P99 (Percentiles)Pairwise evaluationPass@kPlaygroundPromptPrompt (deployed)QQuality gateRRAG (retrieval-augmented generation)RAG evaluationReference-based scoringReference-free scoringRegression testingRelease criteriaRemote evaluationRubricSSafetySampling rateScore distributionScorerSemantic failureService Level Indicator (SLI)Service Level Obj

首次發現: 最近檢查: 內容更新:
文章站內連結

The principles of Braintrust

https://www.braintrust.dev/manifesto

前往原始頁面

Braintrust's principles, beliefs, and practical guidance for building high-quality evals.

語言: en
索引內容摘要

Evals and observability are AI infrastructureThe eval is as important as the model, harness, etc. It should be treated as part of the stack, not an afterthought.If you’re building agents, you need to consider the infrastructure stack the same way you would for any software. What cloud provider, database, what IaC framework? What experiments, which scorers, which rubrics?Agents are now in production, supporting products used by millions. We are past the point of toy demos and neat experiments. Serious products need serious infrastructure. And real infrastructure is observable.

首次發現: 最近檢查: 內容更新:
文章站內連結

Pricing

https://www.braintrust.dev/pricing

前往原始頁面

Start building AI products for free with Braintrust. Transparent pricing for AI evaluation, monitoring, and observability. No hidden fees, pay as you scale.

語言: en
索引內容摘要

ProductResourcesCustomersPricingContact usSign inSign upObserveTrace everythingEvaluateTest what shipsDiscoverFind behaviorsDocsStart buildingBlogInsights and updates .workshops-icon-col { transition: transform 400ms cubic-bezier(0.65, 0, 0.35, 1); } .workshops-icon-playing .workshops-icon-col { transform: translateY(7.5px); } WorkshopsLive sessions .evals-icon-row { transition: transform 350ms ease-in-out; } .evals-icon-playing .evals-icon-row-down { transform: translateY(5px); } .evals-icon-playing .evals-icon-row-up { transform: translateY(-5px); } Eval libraryOpen source researchFoundationsLearn to evalEncyclopediaEval conceptsProductObserveEvaluateDiscoverResourcesDocsBlogWorkshopsEval libraryFoundationsEncyclopediaCustomersPricingContact usSign inSign upPredictable pricing. Designed to scale.Start now. Instrument when you’re ready.No credit card required. Free usage is included for traces, evals, and your whole team.Sign up Start building for freeModel creditsProcessed dataScoresData retentionFeaturesStarterFor everyone$0 / month$10 credits+ tok rates1 GB processed data+ $4/GB10k scores+ $2.50/1k14-day retentionUnlimited users, projects, datasets, playgrounds, and experiments

首次發現: 最近檢查: 內容更新:
文章站內連結

Take the eval maturity assessment

https://www.braintrust.dev/assessment

前往原始頁面

Find gaps in your AI development loop and get an AI quality roadmap.

語言: en
索引內容摘要

ProductResourcesCustomersPricingContact usSign inSign upObserveTrace everythingEvaluateTest what shipsDiscoverFind behaviorsDocsStart buildingBlogInsights and updates .workshops-icon-col { transition: transform 400ms cubic-bezier(0.65, 0, 0.35, 1); } .workshops-icon-playing .workshops-icon-col { transform: translateY(7.5px); } WorkshopsLive sessions .evals-icon-row { transition: transform 350ms ease-in-out; } .evals-icon-playing .evals-icon-row-down { transform: translateY(5px); } .evals-icon-playing .evals-icon-row-up { transform: translateY(-5px); } Eval libraryOpen source researchFoundationsLearn to evalEncyclopediaEval conceptsProductObserveEvaluateDiscoverResourcesDocsBlogWorkshopsEval libraryFoundationsEncyclopediaCustomersPricingContact usSign inSign upEvaluate your AI development loop.Not all agents are created with the same quality and feedback loops. Find out where your team ranks given today's leading agent development workflows.Start assessment01Describe the productName the agent, workflow, or AI feature you want to improve.02Map the loopMark the quality practices your team already has in motion.03Set confidenceShare how confidently you can catch regressions today.React

首次發現: 最近檢查: 內容更新: