ClearTrace Scorecard · Edition 1 of a recurring series
DEX Aggregator Execution Scorecard
Edition 1, July 2026
The rated record of DEX aggregator execution quality: quote accuracy, execution
slippage, revert rates, and RFQ routing at size, measured the same way for every aggregator,
by a third party with no routing product, and frozen so it can't be edited after the fact.
Chain: Ethereum ·
Data frozen: 2026-07-04 04:04 UTC ·
Published: 2026-07-04 ·
Snapshot: scorecard-edition-1.json
SHA-256: a010d40ed952cb8701b96a1bc9c1d00745c3447fd3858b8ef86eca6208d56bb3
Why a rated edition exists
Every aggregator claims best price and high reliability, and those claims are
self-reported, published by the routing product being measured. Meanwhile,
since 1 July 2026, MiCA transposes MiFID-style best-execution duties into
crypto: best-ex must be evidenced, with multi-year record-keeping, not asserted.
Compliance dashboards (e.g. DefiLlama's
MiCA tracker) cover exchange-level obligations, not execution quality at the routing
layer, where the trades actually happen.
This scorecard is the missing instrument: a neutral, methodology-published,
third-party rated record. Each edition is numbered, dated, and hashed. A rating you
can cite in a governance forum, a marketing page, or a best-execution file, and one nobody
(including us) can quietly back-date or revise.
Edition 1 findings
Post-publication note (2026-07-05). A sender-concentration audit run after this
edition froze found that the Odos revert figure below is dominated by address-rotating solver
bots, not user transactions failing. The rate experienced by genuine Odos users on Ethereum
is ≈0.2%. Revert methodology moved to v5 (a stronger sender-level bot filter) on 2026-07-05 and
Edition 2 will use it. Per this scorecard's integrity model, the Edition 1 snapshot, its findings
text, and its SHA-256 hash are unchanged: the record stays frozen, and corrections are
published as dated notes like this one. The other findings are unaffected.
Post-publication note (2026-07-14). The first finding below is withdrawn.
It compares revert rates across venues that do not settle the same way, and that comparison is
not valid. A revert rate measures reliability only where the failing transaction is the
user's own: that is, on a
router, where the user signs and submits the swap.
On a batch-auction venue (CoW Protocol) or an intent venue (1inch Fusion, which settles through
the Limit Order Protocol), a solver or resolver submits instead: an order that cannot be filled
never becomes a transaction at all, so the on-chain revert rate is structurally near zero however
reliably the venue actually fills.
Tokenlon, named above as the reliable end of the spread at 0.04%, is an RFQ venue for which we
never established who submits the settlement transaction. We therefore cannot support the claim
that 0.04% is the rate at which its users' swaps fail, and it should not have been placed
opposite a router. Together with the 2026-07-05 note above,
both ends of the finding are
unsupported: the spread it reports is in substantial part an artifact of how the venues settle,
not a measure of how reliably they fill. The 2026-07-05 note closed by saying the other findings
were unaffected; that was itself incomplete, because this error was already present in the same
finding and we did not catch it.
The caution extends to the revert column of the table below. The figures for CoW Protocol
(0.11%), the 1inch Limit Order Protocol (0.84%), Tokenlon (0.04%), DODO X (0.15%) and Bebop
(8.05%) are not comparable to the router figures beside them and must not be ranked against them.
Bebop runs both taker-submitted RFQ and solver-submitted JAM, so its figure mixes two populations
and is not well defined as a single rate.
ClearTrace now tags every venue with an execution model (router, batch auction, intent, or
unclassified where we have not established who submits) and ranks revert rates only within a
model. Venues we cannot establish are left unclassified and are not ranked at all, rather than
assumed. The
live leaderboard shows each venue's model, and Edition 2
will apply the model guard together with the v5/v6 revert methodology.
Per this scorecard's integrity model, the Edition 1 snapshot, its findings text, and its SHA-256
hash are
unchanged: the record stays frozen, and corrections are published as dated
notes like this one. Finding 2 rests on the same Odos figure addressed in the 2026-07-05 note.
Findings 3 and 4 do not use revert data and are unaffected.
Post-publication note (2026-09-10). Tokenlon's execution-cost figure below is withdrawn.
Most of it was never a cost anyone paid. Dune's Tokenlon spellbook records the
signed order's quoted amount as the amount bought:
tokenlon_v5_ethereum_amm_v1/v2_trades take token_bought_amount_raw from
order.makerAssetAmount on the AMMWrapper Swapped event, while the same
event carries the settleAmount the taker actually received. The amount
sold is read from takerAssetAmount and is correct, because the taker does
pay exactly that. So the error runs one way only, and an execution-cost metric measured against
the quote reads it as a cost.
Decoding every AMM-route fill of a seven-day window (1,310 fills, from transaction receipts)
puts the size of it at a median 47.4 bps: takers received that much more
than the source records, and the quoted amount matches no transfer in the transaction.
About 30 bps of what remains is Tokenlon's own disclosed fee, measured from the same log,
leaving a few basis points of genuine routing cost. Tokenlon's RFQ route reads the same field and
shows no such gap, because an RFQ quote is firm and the quote is the fill.
The caution extends to the execution-cost column of the table below, which is sorted by it.
Tokenlon's position at the bottom of that column is a ranking whatever the prose says, and it is
not supported.
No replacement figure. The window decoded above is a recent one, not this
edition's. The direction is certain for every window — on a route that quotes with
a slippage buffer the recorded quote can only understate what the taker received — so
64.48 bps overstates Tokenlon's execution cost here too. The size of the overstatement in
this edition's window has not been measured, and we are not going to publish a number we did not
measure. It follows that we have also not established whether the count in
“Beating the venue default is rare” is affected: it moves only if the corrected figure falls below
the 10.54 bps baseline, and we have not checked that for this window.
The same column carries a second withdrawn cell, for an unrelated reason.
The 1inch Limit Order Protocol figure (27.67 bps) was withdrawn from ClearTrace's live surfaces on
2026-09-09. There the defect runs the other way: the published metric takes the absolute value of
each fill's gap to the oracle, which folds a 1inch Resolver's own maker-side edge in as if it were
the taker's cost. Over the window then measured that made the venue read roughly twice as good as
the population left after removing the resolver — 13.27 bps as published against
29.73 bps without it. It is named here because a note that corrected one cell of a sorted
column while leaving another withdrawn cell in it would repeat the incompleteness the 2026-07-14
note above had to admit to.
Per this scorecard's integrity model, the Edition 1 snapshot, its findings text, and its SHA-256 hash
are unchanged: the record stays frozen, and corrections are published as dated notes like
this one. Tokenlon's execution cost is no longer served on any live ClearTrace surface, pending a
source that records settled amounts; its revert rate is measured from transaction outcomes, not
from amounts, and is unaffected by this defect.
- Reliability is the widest spread in DeFi routing. Odos reverts
23.78% of routing transactions, about 1 in
4, and 14.8× the direct-Uniswap baseline
(1.61%). At the other end, Tokenlon reverts just
0.04%.
- Accurate quotes don't imply reliable fills. Odos posts a rated
median quote gap of 0 bps (essentially perfect) while
failing 23.78% of its routing transactions on-chain. Quote
accuracy and execution reliability are separate dimensions; a scorecard that reports only
one is marketing.
- RFQ routing at size doesn't close the quote gap. KyberSwap routes
86% of sampled $1M flow through off-chain RFQ desks, yet posts
the widest rated quote gap (2.87 bps).
- Beating the venue default is rare. Against the direct-Uniswap execution
baseline (10.54 bps median), the aggregators that deliver
cheaper realized execution are: Bitget DEX (2.32 bps), 1inch (7.02 bps), Bebop (9.73 bps). Everyone else routes worse than the default
they're competing with.
The rated table
Rated = ≥30 realized fork-simulation quote samples spanning
≥7 days. On-chain metrics (slippage, reverts) cover all
trades/transactions in the window, not a sample. Baseline = the direct-venue default
aggregators are implicitly compared against.
| Aggregator | Status | Exec slippage | Revert rate |
RFQ @ $1M | Quote gap | Quote n | Routing txs |
| Bitget DEX | on-chain only | 2.32 bps | 1.29% | — | — | — | 52,904 |
| 1inch | rated | 7.02 bps | 2.16% | 0% | 0 bps | 968 | 99,483 |
| Bebop | on-chain only | 9.73 bps | 8.05% | — | — | — | 106,185 |
| Uniswap (direct venue) | baseline | 10.54 bps | 1.61% | 0% | 0 bps | 1,001 | 10,774 |
| KyberSwap | rated | 12.08 bps | 4.28% | 86% | 2.87 bps | 920 | 137,934 |
| Odos | rated | 14.3 bps | 23.78% | 0% | 0 bps | 831 | 2,090 |
| CoW Protocol | on-chain only | 15.72 bps | 0.11% | — | — | — | 47,118 |
| OpenOcean | on-chain only | 16.42 bps | 10.95% | — | — | — | 13,892 |
| 1inch Limit Order Protocol | on-chain only | 27.67 bps | 0.84% | — | — | — | 30,111 |
| SushiSwap | on-chain only | 34.9 bps | 4.8% | — | — | — | 26,313 |
| ParaSwap (Velora) | rated | 43.01 bps | 3.83% | 0% | 1 bps | 977 | 25,572 |
| Tokenlon | on-chain only | 64.48 bps | 0.04% | — | — | — | 2,789 |
| DODO X | on-chain only | — | 0.15% | — | — | — | 6,886 |
Verify this edition
The findings above are computed from a frozen snapshot. To verify nothing has changed since
publication, hash the snapshot and compare:
curl -s https://cleartracedata.com/static/data/scorecard-edition-1.json | shasum -a 256
a010d40ed952cb8701b96a1bc9c1d00745c3447fd3858b8ef86eca6208d56bb3
The snapshot is committed to a public git history at
publication time, which independently timestamps it.
Methodology & scope
Full methodology: cleartracedata.com/methodology and the
open-dataset README. In brief: execution slippage is the
median per-fill gap vs a 1-minute VWAP oracle over all on-chain trades; revert rate counts
failed routing transactions from raw on-chain data; quote accuracy is a forward-captured
fork-simulation sample (quoted vs realized); RFQ share is measured from the same sampler at
$1M notional. Aggregators we don't yet quote-sample appear with on-chain metrics only.
Preliminary cells are shown but never rated. Scope: Ethereum for this edition; the live
leaderboard is the free preview layer that accrues between
editions.
Post-publication note (2026-09-25). The sentence above misnames the slippage baseline.
It says execution slippage is measured against “a 1-minute VWAP oracle.” It is not. The query
values both legs of each fill against Dune’s prices.usd at the minute of the fill, and
Dune documents that table as a view of its Coinpaprika feed: exchange-aggregate market data at
5-minute intervals, interpolated to the minute. No volume weighting is applied anywhere in the query.
Every figure on this page was computed exactly as it says elsewhere; only the name of the reference
price was wrong. Logged in our corrections log and registered in our metric-findings registry as
exec-slippage-baseline-is-not-a-vwap-2026-09-25.
Cite or commission
Cite this edition as: ClearTrace DEX Aggregator Execution Scorecard, Edition 1
(July 2026), sha256:a010d40ed952…. Link: https://cleartracedata.com/scorecard/edition-1. For a private cut of your own routing
versus peers (per-pair, per-size, with the failure modes), or a subscription to future rated
editions for a best-execution evidence file,
book a working session.