Benchmarks

Falconer outscores leading AI tools on real support and engineering questions

Falconer's overall win rate · both tests combined
vsNotion
64%
vsAtlassian Rovo
93%
vsClaude Code
54%
vsCodex
68%
Share of decisive judge verdicts won by Falconer · 200 questions · 4 other AI tools on the market · 3 LLM judges
What we tested

Benchmarks for docs and code

We judged performance against 200 real questions from a public support corpus and the codebase of a popular open-source project. Falconer wins head-to-head against all leading tools.

Documentation test

Wix help center

Can a system answer real customer support questions using only a help center? We pointed each tool at the same 6,221 Wix help articles, then scored it on 100 real customer questions from the public WixQA set.

PersonaA Wix customer asking a real support question (billing, domains, site setup) and needs a solution.
Corpus
  • 6,221 articles
  • 100 questions
Models tested
FalconerClaude Opus 4.8 · medium effortNotionClaude Opus 4.8 · effort undisclosedAtlassian Rovomodel undisclosed · Think deeperClaude CodeClaude Opus 4.8 · medium effortCodexGPT-5.5 · medium effort
Code test

Apache Spark

Can a system answer real engineering questions from live source code? We indexed the apache/spark repository and its markdown docs, then scored each tool on 100 Spark questions from Stack Overflow.

PersonaA data engineer who wants working code, an accurate config key, or a root cause.
Corpus
  • apache/spark code package with markdown files
  • 100 questions
Models tested
FalconerClaude Opus 4.8 · high effortNotionClaude Opus 4.8 · effort undisclosedAtlassian Rovomodel undisclosed · Think deeperClaude CodeClaude Opus 4.8 · high effortCodexGPT-5.5 · high effort
The results

Benchmark results

Note: Falconer, Claude Code, and Codex ran at the same thinking-effort level on each test: medium effort on the docs test, high effort on the code test. Notion and Atlassian Rovo do not expose a comparable effort setting.

On real support questions, Falconer wins against all four tools.

Head-to-head win rate

Falconer wins the majority of decisive verdicts in every matchup.

Win share of decisive judge verdicts (6 per question), scored by the weighted-sum rule. Ties excluded; the full record is shown beside each matchup.

vs
Notion
316 wins · 132 losses · 116 ties
70.5%
vs
Atlassian Rovo
503 wins · 66 losses · 31 ties
88.4%
vs
Claude Code
213 wins · 192 losses · 195 ties
52.6%
vs
Codex
314 wins · 186 losses · 100 ties
62.8%

Three scoring methods

Matchups were scored holistically, by weighted formula, and by strict Pareto rule.

We set three different ways to pick each question's winner. Each cell is Falconer's win rate under that rule.

Scoring ruleHow it decidesDefinition of tievsvsvsvs
Weighted-sumused aboveBlend the four axes (0.35 faithfulness + 0.35 helpfulness + 0.20 completeness + 0.10 relevance); a side wins if it leads by more than 0.25 on the 0 to 10 scale.The two scores land within 0.25 of each other.70.5%88.4%52.6%62.8%
HolisticTrust the judge's own overall pick of the better answer.The judge called it a tie.70.5%87.8%55%60.2%
ParetoA side wins only if it scores ≥ on all four axes and > on at least one. The strictest rule.Equal on every quality metric, or mixed (each side better somewhere).74%94.6%51.4%52.5%

Speed

Falconer delivers the fastest full answers: 18.5s median, well ahead of every rival.

Time to full answer and to first token, across the full distribution. Lower and tighter is better.

Time to full answer · Wix help center
← faster · tighter = more consistent05s10s15s20s25s30s35s40sFalconer18.5sNotion27.1sAtlassian Rovo30.4sClaude Code27.3sCodex25.9sp25 to p75 box · whisker to p90 · dot = median
Notion answered 94 of 100 questions; all other systems answered every question.
See exact percentiles
Time to full answer (seconds)
SystemMeanp25Medianp75p90
Falconer19.6s15.3s18.5s23.3s30.5s
Notion27.5s22.3s27.1s31.2s35.8s
Atlassian Rovo31.2s26.5s30.4s34.6s37s
Claude Code29.4s21.1s27.3s36.2s41.9s
Codex27.3s21.1s25.9s31s41.3s
Time to first token · Wix help center
← faster · tighter = more consistent05s10s15s20s25s30s35s40sFalconer2.9sNotion18.3sAtlassian Rovo18.9sClaude Code2sCodex25.7sp25 to p75 box · whisker to p90 · dot = median
Notion answered 94 of 100 questions; all other systems answered every question.
See exact percentiles
Time to first token (seconds)
SystemMeanp25Medianp75p90
Falconer3.3s2.5s2.9s3.8s4.5s
Notion19.5s14s18.3s23.3s30s
Atlassian Rovo18.4s13.8s18.9s21.4s23.7s
Claude Code2.2s1.7s2s2.3s2.9s
Codex27s19.9s25.7s30.8s41.1s

Answer length

Top performers show shorter answer length.

The judges never reward length; it's shown here only as context.

Answer length (characters) · Wix help center
01k2k3k4k5k6k7kFalconer1,741Notion2,367Atlassian Rovo6,032Claude Code1,838Codex1,244p25 to p75 box · whisker to p90 · dot = median
Notion answered 94 of 100 questions; all other systems answered every question.
See exact percentiles
Answer length (characters)
SystemMeanp25Medianp75p90
Falconer1,7681,2421,7412,1382,705
Notion2,4191,8312,3672,8323,944
Atlassian Rovo5,7305,3576,0326,8297,529
Claude Code1,8751,3681,8382,4092,884
Codex1,3049181,2441,5922,091

On real engineering questions, Falconer wins against all four tools.

Head-to-head win rate

Falconer wins the majority of decisive verdicts in every matchup.

Win share of decisive judge verdicts (6 per question), scored by the weighted-sum rule. Ties excluded; the full record is shown beside each matchup.

vs
Notion
244 wins · 179 losses · 177 ties
57.7%
vs
Atlassian Rovo
561 wins · 17 losses · 10 ties
97.1%
vs
Claude Code
215 wins · 168 losses · 217 ties
56.1%
vs
Codex
340 wins · 118 losses · 141 ties
74.2%

Three scoring methods

Matchups were scored holistically, by weighted formula, and by strict Pareto rule.

We set three different ways to pick each question's winner. Each cell is Falconer's win rate under that rule.

Scoring ruleHow it decidesDefinition of tievsvsvsvs
Weighted-sumused aboveBlend the four axes (0.35 faithfulness + 0.35 helpfulness + 0.20 completeness + 0.10 relevance); a side wins if it leads by more than 0.25 on the 0 to 10 scale.The two scores land within 0.25 of each other.57.7%97.1%56.1%74.2%
HolisticTrust the judge's own overall pick of the better answer.The judge called it a tie.58.8%96.2%55%74.2%
ParetoA side wins only if it scores ≥ on all four axes and > on at least one. The strictest rule.Equal on every quality metric, or mixed (each side better somewhere).61.1%97.7%55.7%75.7%

Speed

On the code test, speed is a near tie: every tool but Codex (72s) returns a full answer in 39 to 45s median.

Time to full answer and to first token, across the full distribution. Lower and tighter is better.

Time to full answer · Apache Spark
← faster · tighter = more consistent025s50s75s100s125sFalconer45sNotion40sAtlassian Rovo42.6sClaude Code39.1sCodex72sp25 to p75 box · whisker to p90 · dot = median
Atlassian Rovo answered 98 of 100 questions; all other systems answered every question.
See exact percentiles
Time to full answer (seconds)
SystemMeanp25Medianp75p90
Falconer51.3s31.9s45s64.7s85.4s
Notion45.4s28.6s40s56.2s69.8s
Atlassian Rovo49.3s33.2s42.6s58.3s78.3s
Claude Code45.4s25.3s39.1s55.9s86.1s
Codex78.6s58.7s72s99.8s120.8s
Time to first token · Apache Spark
← faster · tighter = more consistent025s50s75s100s125sFalconer15.5sNotion24.4sAtlassian Rovo12.7sClaude Code1.9sCodex55.7sp25 to p75 box · whisker to p90 · dot = median
Atlassian Rovo answered 98 of 100 questions; all other systems answered every question.
See exact percentiles
Time to first token (seconds)
SystemMeanp25Medianp75p90
Falconer19.1s5.5s15.5s26s42.1s
Notion27.4s12.4s24.4s36.9s50.9s
Atlassian Rovo13.7s11.8s12.7s15.1s17.9s
Claude Code2.1s1.6s1.9s2.2s2.8s
Codex55.1s10.2s55.7s98.5s120.7s

Answer length

Top performers show shorter answer length.

The judges never reward length; it's shown here only as context.

Answer length (characters) · Apache Spark
01k2k3k4k5kFalconer3,205Notion3,655Atlassian Rovo228Claude Code3,036Codex1,802p25 to p75 box · whisker to p90 · dot = median
Atlassian Rovo answered 98 of 100 questions; all other systems answered every question.
See exact percentiles
Answer length (characters)
SystemMeanp25Medianp75p90
Falconer3,2952,5923,2054,0314,764
Notion3,8462,9423,6554,7345,542
Atlassian Rovo5861362284521,676
Claude Code3,0672,4513,0363,7334,411
Codex1,8701,3981,8022,2832,895
Judging details

How the judges decided

The full dataset is public, including every question, reference answer, assistant response, and judge verdict: FalconerAI/falconer-benchmarks.

Real questions used

The questions come from public, third-party sources: real help-center Q&A from WixQA and Apache Spark questions from Stack Overflow. For Spark, we selected questions by user votes and by whether answering them well requires reading the repository. We published every question, the reference answer, and each assistant’s full response.

Judging principles

01

Three frontier judges

Every head-to-head is scored by Claude Opus 4.8, GPT-5.5, and Gemini 3.1 Pro, in both A/B and B/A orderings to neutralize position bias, for six verdicts per question.

02

A reproducible formula

The headline winner is derived from four metric scores with a fixed weighting. Anyone with the per-metric scores gets the same answer, and the lead holds under two other rules too.

03

What counts as a tie

Weighted-sum scores within 0.25 of each other count as ties and are excluded from head-to-head win rates.

04

Pairwise, against a human reference

Every answer is judged head-to-head against Falconer, with the human-written gold answer (the WixQA reference, or the accepted Stack Overflow answer) as the reference both sides are measured against, rather than grading either answer in isolation.

05

No web access for any agent

Every system runs sealed to its corpus. Web tools are disabled and verified at runtime, so results reflect retrieval and grounding.

06

Public, reproducible corpora

Both question sets are public (the WixQA support corpus and the apache/spark codebase with its Stack Overflow Q&A), at 100 questions each, so anyone can rerun the same benchmark against their own system.

Notes
  • The displayed records use the headline weighted-sum rule: a verdict is a tie when the two weighted scores land within 0.25 of each other on the 0 to 10 scale, too close for the judge to separate (the scoring table shows how each of the three rules defines a tie). Head-to-head % counts wins among the decisive (non-tie) verdicts, the standard way win rates are reported.
  • Each question set is 100 items drawn from a public benchmark; larger samples tighten the confidence intervals.
  • Coverage varies slightly. Notion answered 94 of 100 support questions; on the other 6 it returned an interactive clarification form instead of an answer, so those are excluded. Rovo answered 98 of 100 engineering questions. Win/loss/tie counts are tallied over the questions each system actually answered (6 verdicts per answered question), so they total below 600 for those systems. One Codex verdict on the engineering test was an unparseable judge response and is also excluded, so that record totals 599.
  • On the engineering test, Atlassian Rovo is a structural baseline rather than a like-for-like rival: the apache/spark repository was hosted in Bitbucket, which Rovo barely reads, so its short answers reflect missing source access rather than answer quality. Read that matchup as retrieval-with-code vs retrieval-without-code; Notion, Claude Code, and Codex are the code-capable cohort.
  • Per-metric averages and head-to-head wins can diverge: a system can post a higher average on one metric yet win fewer questions, because each question’s winner is decided from all four metrics together.
  • Ties (weighted-sum scores within 0.25 of each other) are excluded from head-to-head %. Tie rates ranged from ~2% (Rovo) to ~36% (Claude Code); closer matchups naturally produce more ties.
  • The Stack Overflow reference is one valid solution, not the only correct answer; judges credit alternate-but-correct approaches.
  • Falconer, Claude Code, and Codex run named models at the same thinking-effort level on each test: medium effort on the support test, high effort on the engineering test. Notion exposes its model (Opus 4.8) but not its effort level; Rovo ran in its "Think deeper" mode on an undisclosed model.
  • Results are a point-in-time snapshot from a June 2026 run.
How we measure quality

Four metrics, deliberately weighted

Every answer scored 0 to 10 against the human reference.

Faithfulness35%
correctness

Every claim is correct and supported by the source, with no fabricated steps, wrong API names, invented config keys, or made-up numbers, even when the rest of the answer reads well.

Is every claim true to the source?

Helpfulness35%
actionability

Concrete and actionable: real UI paths, working code, specific steps and version numbers. Vague guidance and "check the docs" deflections are penalized.

Can the reader act on it directly?

Completeness20%
recall

Covers every critical step or fact needed to actually resolve the question. Tangential background does not raise this score.

Did it cover everything needed to finish the job?

Relevance10%
precision

The answer is relevant. Sometimes an answer can be factually true, but irrelevant to the problem in question. Tangents, unrequested background, and citation dumps are penalized. Length itself is not.

How much is on-target signal vs. noise?

The weightingscore = 0.35·faithfulness + 0.35·helpfulness + 0.20·completeness + 0.10·relevance
  • Correct and actionable matter most, so faithfulness and helpfulness get 35% each. A fluent answer with the wrong API name is worthless.
  • Completeness (20%) catches answers that skip a critical step.
  • Relevance (10%) penalizes off-topic padding.
  • Completeness and relevance deliberately pull against each other: adding content lifts completeness but costs relevance, so neither rewards bloat.

The real benchmark is your own data

These results come from publicly available data. Falconer is even better when connected to your own code, docs, and sources like Linear, Slack, and meeting notes.

Get startedRequest demo