chore: import upstream snapshot with attribution
OSV-Scanner (Scheduled) / scan-scheduled (push) Failing after 0s
Create Release / test-gate (push) Has been cancelled
Create Release / release-gate (push) Has been cancelled
Create Release / ci-gate (push) Has been cancelled
Create Release / version-check (push) Has been cancelled
Create Release / e2e-test-gate (push) Has been cancelled
Create Release / responsive-test-gate (push) Has been cancelled
Create Release / compat-test-gate (push) Has been cancelled
Create Release / compose-integration-gate (push) Has been cancelled
Create Release / vulture-gate (push) Has been cancelled
Create Release / build (push) Has been cancelled
Create Release / provenance (push) Has been cancelled
Create Release / prerelease-docker (push) Has been cancelled
Create Release / publish-docker (push) Has been cancelled
Create Release / create-release (push) Has been cancelled
Create Release / cleanup-changelog (push) Has been cancelled
Create Release / trigger-pypi (push) Has been cancelled
Create Release / monitor-pypi (push) Has been cancelled
Create Release / Clean up orphan prerelease tags and signatures (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [research-form] (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [research-metrics] (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [research-workflow] (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [settings-core] (push) Has been cancelled
CodeQL Advanced / Analyze (javascript-typescript) (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [history-news] (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [library] (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [link-analytics] (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [chat-core] (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [chat-lifecycle] (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [error-benchmark] (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [settings-pages] (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) (push) Has been cancelled
Docker Tests (Consolidated) / Accessibility Tests (push) Has been cancelled
Docker Tests (Consolidated) / LLM Unit Tests (push) Has been cancelled
Docker Tests (Consolidated) / LLM Example Tests (push) Has been cancelled
Docker Tests (Consolidated) / Production Image Smoke Test (push) Has been cancelled
Docker Tests (Consolidated) / Infrastructure Tests (push) Has been cancelled
OSSF Scorecard / OSSF Security Scorecard Analysis (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [mobile] (push) Has been cancelled
Backwards Compatibility / Verify Encryption Constants (push) Has been cancelled
Backwards Compatibility / PyPI Version Compatibility (push) Has been cancelled
Backwards Compatibility / Database Migration Tests (push) Has been cancelled
CodeQL Advanced / Analyze (python) (push) Has been cancelled
Docker Tests (Consolidated) / detect-changes (push) Has been cancelled
Docker Tests (Consolidated) / Build Test Image (push) Has been cancelled
Docker Tests (Consolidated) / All Pytest Tests + Coverage (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [accessibility] (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [api-crud] (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [auth-login] (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [auth-pages] (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [auth-register] (push) Has been cancelled
OSV-Scanner (Scheduled) / scan-scheduled (push) Failing after 0s
Create Release / test-gate (push) Has been cancelled
Create Release / release-gate (push) Has been cancelled
Create Release / ci-gate (push) Has been cancelled
Create Release / version-check (push) Has been cancelled
Create Release / e2e-test-gate (push) Has been cancelled
Create Release / responsive-test-gate (push) Has been cancelled
Create Release / compat-test-gate (push) Has been cancelled
Create Release / compose-integration-gate (push) Has been cancelled
Create Release / vulture-gate (push) Has been cancelled
Create Release / build (push) Has been cancelled
Create Release / provenance (push) Has been cancelled
Create Release / prerelease-docker (push) Has been cancelled
Create Release / publish-docker (push) Has been cancelled
Create Release / create-release (push) Has been cancelled
Create Release / cleanup-changelog (push) Has been cancelled
Create Release / trigger-pypi (push) Has been cancelled
Create Release / monitor-pypi (push) Has been cancelled
Create Release / Clean up orphan prerelease tags and signatures (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [research-form] (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [research-metrics] (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [research-workflow] (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [settings-core] (push) Has been cancelled
CodeQL Advanced / Analyze (javascript-typescript) (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [history-news] (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [library] (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [link-analytics] (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [chat-core] (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [chat-lifecycle] (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [error-benchmark] (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [settings-pages] (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) (push) Has been cancelled
Docker Tests (Consolidated) / Accessibility Tests (push) Has been cancelled
Docker Tests (Consolidated) / LLM Unit Tests (push) Has been cancelled
Docker Tests (Consolidated) / LLM Example Tests (push) Has been cancelled
Docker Tests (Consolidated) / Production Image Smoke Test (push) Has been cancelled
Docker Tests (Consolidated) / Infrastructure Tests (push) Has been cancelled
OSSF Scorecard / OSSF Security Scorecard Analysis (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [mobile] (push) Has been cancelled
Backwards Compatibility / Verify Encryption Constants (push) Has been cancelled
Backwards Compatibility / PyPI Version Compatibility (push) Has been cancelled
Backwards Compatibility / Database Migration Tests (push) Has been cancelled
CodeQL Advanced / Analyze (python) (push) Has been cancelled
Docker Tests (Consolidated) / detect-changes (push) Has been cancelled
Docker Tests (Consolidated) / Build Test Image (push) Has been cancelled
Docker Tests (Consolidated) / All Pytest Tests + Coverage (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [accessibility] (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [api-crud] (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [auth-login] (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [auth-pages] (push) Has been cancelled
Docker Tests (Consolidated) / UI Tests (Puppeteer) [auth-register] (push) Has been cancelled
This commit is contained in:
@@ -0,0 +1,194 @@
|
||||
#!/usr/bin/env python
|
||||
"""
|
||||
Example script for running benchmarks using the Local Deep Research benchmarking framework.
|
||||
|
||||
This script demonstrates how to run SimpleQA and BrowseComp benchmarks programmatically.
|
||||
"""
|
||||
|
||||
import argparse
|
||||
import os
|
||||
|
||||
|
||||
from local_deep_research.api.benchmark_functions import (
|
||||
compare_configurations,
|
||||
evaluate_browsecomp,
|
||||
evaluate_simpleqa,
|
||||
)
|
||||
|
||||
|
||||
def main():
|
||||
"""Run benchmark examples."""
|
||||
parser = argparse.ArgumentParser(description="LDR Benchmark Examples")
|
||||
parser.add_argument(
|
||||
"--benchmark",
|
||||
choices=["simpleqa", "browsecomp", "compare"],
|
||||
default="simpleqa",
|
||||
help="Benchmark to run",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--examples", type=int, default=10, help="Number of examples to use"
|
||||
)
|
||||
parser.add_argument(
|
||||
"--iterations", type=int, default=3, help="Number of search iterations"
|
||||
)
|
||||
parser.add_argument(
|
||||
"--questions", type=int, default=3, help="Questions per iteration"
|
||||
)
|
||||
parser.add_argument(
|
||||
"--search-tool", default="searxng", help="Search tool to use"
|
||||
)
|
||||
parser.add_argument(
|
||||
"--human-eval", action="store_true", help="Use human evaluation"
|
||||
)
|
||||
parser.add_argument(
|
||||
"--output-dir",
|
||||
default="benchmark_results",
|
||||
help="Directory to save results",
|
||||
)
|
||||
|
||||
args = parser.parse_args()
|
||||
|
||||
# Create output directory if it doesn't exist
|
||||
os.makedirs(args.output_dir, exist_ok=True)
|
||||
|
||||
print(f"Running {args.benchmark} benchmark with {args.examples} examples")
|
||||
|
||||
# Run the specified benchmark
|
||||
if args.benchmark == "simpleqa":
|
||||
run_simpleqa_example(args)
|
||||
elif args.benchmark == "browsecomp":
|
||||
run_browsecomp_example(args)
|
||||
elif args.benchmark == "compare":
|
||||
run_comparison_example(args)
|
||||
else:
|
||||
print(f"Unknown benchmark: {args.benchmark}")
|
||||
|
||||
|
||||
def run_simpleqa_example(args):
|
||||
"""Run SimpleQA benchmark."""
|
||||
print("\n=== SimpleQA Benchmark ===")
|
||||
print(f"Running with {args.examples} examples")
|
||||
print(f"Search iterations: {args.iterations}")
|
||||
print(f"Questions per iteration: {args.questions}")
|
||||
print(f"Search tool: {args.search_tool}")
|
||||
print(f"Human evaluation: {args.human_eval}")
|
||||
print(f"Output directory: {args.output_dir}")
|
||||
print("=" * 30)
|
||||
|
||||
# Run benchmark
|
||||
result = evaluate_simpleqa(
|
||||
num_examples=args.examples,
|
||||
search_iterations=args.iterations,
|
||||
questions_per_iteration=args.questions,
|
||||
search_tool=args.search_tool,
|
||||
human_evaluation=args.human_eval,
|
||||
output_dir=args.output_dir,
|
||||
)
|
||||
|
||||
# Print results
|
||||
if "metrics" in result:
|
||||
print("\nResults:")
|
||||
print(f" Accuracy: {result['metrics'].get('accuracy', 0):.3f}")
|
||||
print(f" Total examples: {result['total_examples']}")
|
||||
print(f" Correct answers: {result['metrics'].get('correct', 0)}")
|
||||
print(
|
||||
f" Average time: {result['metrics'].get('average_processing_time', 0):.2f}s"
|
||||
)
|
||||
print(f"\nReport saved to: {result.get('report_path', 'N/A')}")
|
||||
else:
|
||||
print("\nBenchmark completed without evaluation")
|
||||
print(f" Results saved to: {result.get('results_path', 'N/A')}")
|
||||
|
||||
|
||||
def run_browsecomp_example(args):
|
||||
"""Run BrowseComp benchmark."""
|
||||
print("\n=== BrowseComp Benchmark ===")
|
||||
print(f"Running with {args.examples} examples")
|
||||
print(f"Search iterations: {args.iterations}")
|
||||
print(f"Questions per iteration: {args.questions}")
|
||||
print(f"Search tool: {args.search_tool}")
|
||||
print(f"Human evaluation: {args.human_eval}")
|
||||
print(f"Output directory: {args.output_dir}")
|
||||
print("=" * 30)
|
||||
|
||||
# Run benchmark
|
||||
result = evaluate_browsecomp(
|
||||
num_examples=args.examples,
|
||||
search_iterations=args.iterations,
|
||||
questions_per_iteration=args.questions,
|
||||
search_tool=args.search_tool,
|
||||
human_evaluation=args.human_eval,
|
||||
output_dir=args.output_dir,
|
||||
)
|
||||
|
||||
# Print results
|
||||
if "metrics" in result:
|
||||
print("\nResults:")
|
||||
print(f" Accuracy: {result['metrics'].get('accuracy', 0):.3f}")
|
||||
print(f" Total examples: {result['total_examples']}")
|
||||
print(f" Correct answers: {result['metrics'].get('correct', 0)}")
|
||||
print(
|
||||
f" Average time: {result['metrics'].get('average_processing_time', 0):.2f}s"
|
||||
)
|
||||
print(f"\nReport saved to: {result.get('report_path', 'N/A')}")
|
||||
else:
|
||||
print("\nBenchmark completed without evaluation")
|
||||
print(f" Results saved to: {result.get('results_path', 'N/A')}")
|
||||
|
||||
|
||||
def run_comparison_example(args):
|
||||
"""Run configuration comparison."""
|
||||
print("\n=== Configuration Comparison ===")
|
||||
print(f"Dataset: {args.benchmark}")
|
||||
print(f"Examples per configuration: {args.examples}")
|
||||
print(f"Output directory: {args.output_dir}")
|
||||
print("=" * 30)
|
||||
|
||||
# Define configurations to compare
|
||||
configurations = [
|
||||
{
|
||||
"name": "Base Config",
|
||||
"search_tool": args.search_tool,
|
||||
"iterations": 1,
|
||||
"questions_per_iteration": 3,
|
||||
},
|
||||
{
|
||||
"name": "More Iterations",
|
||||
"search_tool": args.search_tool,
|
||||
"iterations": 3,
|
||||
"questions_per_iteration": 3,
|
||||
},
|
||||
{
|
||||
"name": "More Questions",
|
||||
"search_tool": args.search_tool,
|
||||
"iterations": 1,
|
||||
"questions_per_iteration": 5,
|
||||
},
|
||||
]
|
||||
|
||||
# Run comparison
|
||||
result = compare_configurations(
|
||||
dataset_type="simpleqa", # Use SimpleQA for faster comparison
|
||||
num_examples=args.examples,
|
||||
configurations=configurations,
|
||||
output_dir=args.output_dir,
|
||||
)
|
||||
|
||||
# Print results
|
||||
print("\nComparison Results:")
|
||||
print(f" Configurations tested: {result['configurations_tested']}")
|
||||
print(f" Report saved to: {result['report_path']}")
|
||||
|
||||
# Print brief comparison table
|
||||
print("\nResults Summary:")
|
||||
print("Configuration | Accuracy | Avg. Time")
|
||||
print("--------------- | -------- | ---------")
|
||||
for res in result["results"]:
|
||||
name = res["configuration_name"]
|
||||
acc = res.get("metrics", {}).get("accuracy", 0)
|
||||
time = res.get("metrics", {}).get("average_processing_time", 0)
|
||||
print(f"{name:15} | {acc:.3f} | {time:.2f}s")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
Reference in New Issue
Block a user