Testing Methodology

How BunusRadar Tests AI Tools

Every AI tool reviewed on BunusRadar is tested using a standardized, reproducible evaluation framework. We define our criteria before testing, run identical workloads across competing tools, and publish our methodology so you can evaluate our conclusions.

Last updated: October 2026 · Maintained by Alex Vance, AI Architecture Lead

Why We Publish Our Methodology

Most technology reviews don't tell you how they reached their conclusions. A reviewer says "Tool A is better than Tool B" without explaining the workloads tested, the scoring criteria, or the conditions. That makes it impossible to know whether the comparison is valid for your use case.

At BunusRadar, we publish our evaluation framework so you can assess our methodology the same way you'd assess a research paper: examine the criteria, check whether our test conditions match your scenario, and form your own conclusions.

The 5 Evaluation Dimensions

Each AI tool is scored across these five dimensions. Weights are category-specific — see individual benchmark reports for category-adjusted weights.

Task Accuracy

Weight: 35%

The percentage of tasks completed correctly without manual correction. We define "correct" with explicit pass/fail criteria before testing begins — never retroactively.

For coding tools: does the code run, does it produce correct output, does it follow the specified requirements? For chatbots: factual accuracy verified against primary sources. For image generators: prompt adherence scored on a 5-point rubric.

Speed & Latency

Weight: 20%

Median time-to-first-token and total task completion time, measured across 3 runs and averaged. Network latency is normalized.

Tested on a standardized fiber connection (500 Mbps symmetric). Outliers (>2σ from median) are noted but excluded from the headline figure.

Cost Per Task

Weight: 20%

Real-world cost to complete one standard task at the tool's listed pricing tier. Calculated as: (tokens consumed × price per token) + subscription cost amortized over task volume.

We report cost in USD at public pricing. Enterprise/negotiated pricing is noted separately. Pricing snapshots are dated — check the article date.

Hallucination Rate

Weight: 15%

For AI tools: the percentage of outputs that contain factually incorrect, fabricated, or confidently wrong information. Verified against primary sources.

We use a structured evaluation rubric: 0 (no hallucinations), 1 (minor inaccuracy), 2 (significant error), 3 (fabricated content). Overall score is the mean across the task suite.

User Experience

Weight: 10%

Qualitative assessment of interface design, onboarding friction, documentation quality, and overall workflow integration. Rated 1–10 by two independent reviewers.

UX is intentionally weighted lower than performance metrics. We prefer tools that are harder to learn but more powerful, unless the category targets non-technical users.

Standard Test Environment

BunusRadar Test Rig — Spec Sheet
Hardware: MacBook Pro M3 Max, 36GB
OS: macOS Sequoia 15.x
Network: 500 Mbps symmetric fiber
Browser: Chrome stable (latest)
API Testing: Cursor + direct API calls
Pricing Snapshot: Public tier, dated
Run Count: 3× per task, median reported
Blind Review: Second analyst validates

Conflict of Interest Policy

  • We do not accept payment for positive reviews. Editorial conclusions are not for sale.
  • Affiliate links, where present, are disclosed at the article level. They do not influence test scores.
  • Tools provided by vendors for review are treated identically to tools purchased independently.
  • If a vendor disputes our findings, we publish their response alongside our original data.
  • BunusRadar has no investor or ownership relationships with any tool we review.

See the Methodology in Action

The BunusRadar Research Hub publishes our full benchmark datasets — including raw scores, per-task breakdowns, and reproducible results for every AI tool category we test.

View Research & Benchmarks