ClearTrace Scorecard · Edition 3 of a recurring series
DEX Aggregator Execution Scorecard
Edition 3, September 2026
The rated record of DEX aggregator execution quality: quote accuracy, execution
slippage, and revert rates, measured the same way for every aggregator on this board,
by a third party with no routing product, and frozen so the numbers can't be edited after
the fact. The board is not a complete census: venues withheld from ClearTrace's published
outputs, and venues we do not yet sample, do not appear on it at all.
Chain: Ethereum ·
Snapshot frozen: 2026-09-08 17:03 UTC ·
On-chain data: 7-day window to 2026-09-08 ·
Published: 2026-09-08 ·
Snapshot: scorecard-edition-3.json
SHA-256: 0c36db6ecc481a1d60ed4170d79b3f4d7f4c6b385ba399742c05617de15f118c
Why a rated edition exists
Every aggregator claims best price and high reliability, and those claims are
self-reported, published by the routing product being measured. Meanwhile MiCA
transposes MiFID-style best-execution duties into crypto: under
Regulation (EU)
2023/1114, Article 78 obliges a CASP executing client orders to take all necessary steps
to obtain the best result and to evidence it on request, with record-keeping under
Article 68. The regime applied from 30 December 2024, and the transitional window that let
existing providers keep operating without full authorisation closed, at the latest,
on 1 July 2026, so evidencing best execution is already in force.
Compliance dashboards (e.g. DefiLlama's
MiCA tracker) cover exchange-level obligations, not execution quality at the routing
layer, where the trades actually happen.
This scorecard is a neutral, methodology-published, third-party rated
record: we operate no router and take no order flow, the method is published, and
each edition is numbered, dated and hashed. A rating you can cite in a governance forum, a
marketing page, or a best-execution file. The snapshot below is what the hash
covers, so the numbers cannot be back-dated or quietly revised; the findings text on this
page is not hashed, and corrections to it ship as dated, additive notes whose full history
is public in git.
Edition 3 findings
Post-publication note (2026-09-08). One sentence in the first finding is withdrawn.
The first finding below says the 5.2× headline gap between OKX and SushiSwap “is substantially a
difference in traffic mix.” That names a cause, and this scorecard’s own rule for findings
is that they state only what is computable from the frozen snapshot: a count, a span, a rate, a
decomposition, never a cause. The count that sentence rests on is already stated beside it — the
residual term exceeds half the headline gap in 31 of 36 venue pairs — and that is the claim we
stand behind. The finding’s own closing caveat says the residual bucket sits inside a corridor set
by our classifier’s thresholds, which is exactly why it cannot carry a causal reading. Withdrawn:
the “traffic mix” sentence. Retained: every number in the finding. Added for the record: the
opening claim that most of what a headline revert rate counts is not the user’s own failed swap is
supported by a count the page did not print — residual reverted transactions exceed genuine-user
reverted transactions in 9 of 9 routers in this snapshot. Per this scorecard’s integrity
model the snapshot, the findings text and the SHA-256 hash are unchanged; corrections are
published as dated notes like this one.
Post-publication note (2026-09-08). OpenOcean’s quote-accuracy cell should read
stale, not rated. The rated table below shows OpenOcean at 10.23 bps over
6,983 realized quote samples with a rated badge. The figure is real. The badge is not: the last
successful OpenOcean quote sample on Ethereum was captured 2026-08-20 19:05 UTC, and every request since
— 1,716 of them, at roughly 104 a day — returned no quote, because the venue’s endpoint
began answering our sampler with a bot challenge on 21 August. The table’s printed rule for
rated (≥30 realized samples spanning ≥7 days) is met on frozen history; the recency rule we
apply in code (newest successful sample no older than 3 days) is failed by 16 days. That recency rule
landed the same day this edition froze, and the edition was built from a data export eight minutes
older than the first one to carry the corrected badge. A cell that has gone quiet for 19 days is not a
rated cell. Read the 10.23 bps as OpenOcean’s last measured value, not a current one, and read the
“widest quote gap on the board” as a statement about August. Per this scorecard’s
integrity model the snapshot, the findings text and the SHA-256 hash are unchanged; Edition 4
carries this correction forward in its errata.
Post-publication note (2026-09-24). Every quote-gap figure below was computed over four
chains, not Ethereum. Four of them are restated here. This edition is scoped to Ethereum
and its revert and execution-cost columns are. Its quote-gap column is not: the code that built it
filtered the on-chain half to one chain and never filtered the quote samples at all, so each gap
pooled Ethereum, Base, Arbitrum and Optimism fills. Nothing in the output showed it, because the
four sampled pairs carry the same labels on every chain. At this edition’s freeze
39.3% of the 110,676 samples behind these figures were not Ethereum. Rebuilt from the same
corpus at the same freeze, scoped to Ethereum alone, the affected cells read:
OpenOcean 11.06 bps (published 10.23), OKX 3.67 (published 2.50),
LI.FI 1.33 (published 0.71), KyberSwap 0.78 (published 0.26).
Every one moves the same way, because L2 fills quote tighter, so each published figure flattered
the venue against its own Ethereum samples. 1inch, Bebop, Uniswap, Sushi, ParaSwap and Odos are
unchanged at their published values. No verdict in this edition changes: OpenOcean is the widest
gap on either basis, and no venue crosses the 2.0 bps threshold that would add or remove a
quote-accuracy finding. The restatement is trustworthy because the rebuild reproduces this
edition’s published pooled figures exactly before the chain filter is applied, venue
for venue. Per this scorecard’s integrity model the snapshot, the findings text and the
SHA-256 hash are unchanged; the prose above still prints the pooled figures, and the four
values in this note supersede them.
Post-publication note (2026-09-10). Tokenlon's execution-cost figure below is withdrawn.
Most of it was never a cost anyone paid. Dune's Tokenlon spellbook records the
signed order's quoted amount as the amount bought:
tokenlon_v5_ethereum_amm_v1/v2_trades take
token_bought_amount_raw from
order.makerAssetAmount on the AMMWrapper
Swapped event, while the same
event carries the
settleAmount the taker actually received. The amount
sold is read from
takerAssetAmount and is correct, because the taker does
pay exactly that. So the error runs one way only, and an execution-cost metric measured against
the quote reads it as a cost.
Decoding every AMM-route fill of a seven-day window (1,310 fills, from transaction receipts)
puts the size of it at a median
47.4 bps: takers received that much more
than the source records, and the quoted amount matches
no transfer in the transaction.
About 30 bps of what remains is Tokenlon's own disclosed fee, measured from the same log,
leaving a few basis points of genuine routing cost. Tokenlon's RFQ route reads the same field and
shows no such gap, because an RFQ quote is firm and the quote
is the fill.
The caution extends to the execution-cost column of the table below, which is
sorted by it.
Tokenlon's position at the bottom of that column is a ranking whatever the prose says, and it is
not supported.
What this edition would have printed. The seven-day window decoded above
(2026-09-02 to 09-09) overlaps this edition's own execution window (2026-09-01 to 09-08) on six
of seven days and reproduces a near-identical published figure, so the correction can be stated
here rather than only gestured at: measured against what takers received, Tokenlon reads about
25 bps, not the 65.73 bps printed below. That is still above the
10.93 bps direct-Uniswap baseline, so the count in
“Most aggregators do not beat
the venue default” — 5 of 11 cheaper, the other 6 worse — does not change.
The magnitude does, and Tokenlon should not be cited as an outlier on this axis.
The same column carries a second withdrawn cell, for an unrelated reason.
The 1inch Limit Order Protocol figure (11.71 bps) was withdrawn from ClearTrace's live surfaces on
2026-09-09. There the defect runs the other way: the published metric takes the absolute value of
each fill's gap to the oracle, which folds a 1inch Resolver's own maker-side edge in as if it were
the taker's cost. Over the window then measured that made the venue read roughly twice as good as
the population left after removing the resolver — 13.27 bps as published against
29.73 bps without it. It is named here because a note that corrected one cell of a sorted
column while leaving another withdrawn cell in it would repeat the incompleteness
Edition 1's 2026-07-14 note had to admit to.
Per this scorecard's integrity model, the Edition 3 snapshot, its findings text, and its SHA-256 hash
are
unchanged: the record stays frozen, and corrections are published as dated notes like
this one. Tokenlon's execution cost is no longer served on any live ClearTrace surface, pending a
source that records settled amounts; its revert rate is measured from transaction outcomes, not
from amounts, and is unaffected by this defect.
Post-publication note (2026-10-01). The findings below describe the v6 user-only revert rate as a rate among “genuine users”
counted by “our v6 classifier”. It is neither. The v6 predicate is a threshold on each sender’s own outcomes, not a classification of who the sender is: it keeps senders
whose transactions to the venue succeeded at least 90% of the time over the window (personal revert rate below 10%, with at
least one success) and drops the rest. Two consequences follow. The figure is bounded under 10% by construction, and a sender
whose few transactions included one failure is removed from the numerator and the denominator alike. Read every
“genuine-user” figure on this page as the failure rate among consistently-succeeding senders, not as the rate
people experience and not as the share of all attempts that fail. The residual term the first finding decomposes is
likewise everything the threshold removed, which includes searcher traffic but is not a measurement of it. The frozen
snapshot’s own
note field carries the same wording (“senders classified as genuine users”,
“the rate users actually experience”) and is left as it is, because it is inside the hash. No number on this page
changes. Per this scorecard’s integrity model the snapshot, the findings text and the SHA-256 hash are
unchanged; Edition 4 carries this correction forward in its errata, and the filter is set out in full on the
methodology page.
- Most of what a headline revert rate counts is not the user's own failed swap. Every headline here splits into two terms: transactions from senders our v6 classifier counts as genuine users (personal revert below 10%), and everything else it did not classify as a bot. The second term is the larger part of the difference between venues: across the 9 pooled routers it accounts for more than half the headline gap in 31 of 36 venue pairs. So the 5.2× headline gap between OKX (5.02%) and SushiSwap (0.96%) is substantially a difference in traffic mix; on the user-only rate the same two are 1.03% and 0.3%, a 3.4× gap. What this does not show: the residual bucket is bounded by the classifier's own thresholds, so its pooled rate (19.2%–28.0% here) sits inside a corridor we set, and should not be read as a measured constant. Only the share of traffic in it (2.8%–15.9%) varies freely. Every reliability finding below is stated on the user-only rate.
- Genuine-user revert rates span 0.80 percentage points; headline rates span 4.06. Genuine-user revert rates (v6) across the 8 routers where the rate is comparable and uncontaminated run from 0.23% (1inch) to 1.03% (OKX). 7 of the 8 route measurably more reliably for genuine users than the direct-Uniswap default (1.04%). OKX is lower but within measurement error (two-proportion z, p > 0.05), so it should be read as no worse rather than better. The baseline itself, and the 1 router with no genuine-user rate this edition (Odos), are outside this pool.
- Reverts are not comparable across execution models. A revert rate is a reliability signal only where the failing transaction is the user's own: on a router, where the user submits the swap. On a batch auction (CoW Protocol) or an intent venue (1inch Fusion), a solver or resolver submits, so an order that cannot be filled never becomes a transaction and the on-chain revert rate is structurally near zero however reliably the venue fills. The table tags every venue with its model; the reliability findings above rank routers only, and non-router rates are shown flagged n/c, never ranked.
- Most aggregators do not beat the venue default. Against the direct-Uniswap execution baseline (10.93 bps median), 5 of the 11 venues we measure on-chain deliver cheaper realized execution: Bitget DEX (4.14 bps), 1inch (7.45 bps), OpenOcean (9.34 bps), CoW Protocol (9.66 bps), Bebop (10.74 bps). The other 6 route worse than the default they are competing with. 3 venues carry no on-chain execution data this edition (Odos, LI.FI, DODO X) and are unmeasured here. Unlike revert rate, realized execution cost is a per-fill outcome the trader gets whoever submits the transaction, so it is compared across every settlement model; the model is shown for each venue so a reader can scope it further.
Provenance: Bitget DEX and Bebop are named above, and we have not ingested their own published deployment addresses. Their figures rest on our labelling of which contracts belong to them, and should be read as provisional.
The rated table
Rated = ≥30 realized fork-simulation quote samples spanning
≥7 days. On-chain metrics (slippage, reverts) cover every
trade/transaction we attribute to the venue in the window, not a sampled subset of them,
but attribution is entrypoint-framed, so the denominator is the flow entering
through contracts we can tie to the venue, not the venue's total volume. Baseline = the direct-venue default
aggregators are implicitly compared against. Revert rate shows the headline with the v6
genuine-user rate in parentheses; n/c marks a venue whose settlement model
(batch auction, intent, or unclassified where we have not established who submits the
settlement transaction) makes its on-chain revert rate not comparable to a router's.
A venue marked wound down has shut down: its numbers are frozen history for the
window measured, not a live venue. A venue marked unestablished is one whose own
published deployment addresses we have not ingested: its numbers rest on our labelling
rather than on the venue's list, and should be read as provisional. An execution-cost cell
marked all-chain means the venue has no row on this chain, so the figure is its
median across every chain we measure, shown as context and excluded from every finding.
A dash (—) means we hold no measurement for that cell this edition, whether the venue
had no qualifying traffic, fell below the query's reporting floor, or is not yet
quote-sampled; it never means zero. Router counts differ between findings because each
states its own pool: the residual decomposition includes the venue default, the
genuine-user comparison does not.
| Aggregator | Status | Model | Exec slippage | Revert rate |
Quote gap | Quote n | Routing txs |
| Bitget DEX · unestablished | on-chain only | Router | 4.14 bps | 1.72% (0.69% users) | — | — | 35,895 |
| 1inch | rated | Router | 7.45 bps | 1.16% (0.23% users) | 0 bps | 12,021 | 126,629 |
| OpenOcean | rated | Router | 9.34 bps | 2.16% (0.32% users) | 10.23 bps | 6,983 | 8,410 |
| CoW Protocol | on-chain only | Batch auction | 9.66 bps | 0.22% n/c | — | — | 27,258 |
| Bebop · unestablished | rated | Unclassified | 10.74 bps | 2.61% n/c | 0 bps | 8,512 | 18,952 |
| Uniswap (direct venue) | baseline | Router | 10.93 bps | 2.61% (1.04% users) | 0 bps | 11,503 | 145,635 |
| KyberSwap | rated | Router | 11.48 bps | 1.26% (0.62% users) | 0.26 bps | 11,560 | 294,054 |
| 1inch Limit Order Protocol | on-chain only | Intent | 11.71 bps | 0.48% n/c | — | — | 47,754 |
| SushiSwap | rated | Router | 11.8 bps | 0.96% (0.3% users) | 0 bps | 11,081 | 53,299 |
| ParaSwap (Velora) | rated | Router | 19.52 bps | 1.58% (0.34% users) | 1 bps | 11,585 | 33,133 |
| OKX | rated | Router | 20.5 bps | 5.02% (1.03% users) | 2.5 bps | 9,875 | 39,167 |
| Tokenlon · unestablished | on-chain only | Unclassified | 65.73 bps | 0% n/c | — | — | 2,461 |
| Odos · Wound down 2026-07-30 · data through Jul 2026 | rated | Router | — | — | 0 bps | 3,109 | — |
| LI.FI | rated | Router | — | 2.36% (0.43% users) | 0.71 bps | 8,966 | 45,584 |
| DODO X | on-chain only | Unclassified | — | 0.12% n/c | — | — | 9,625 |
Verify this edition
The findings above are computed from a frozen snapshot. To verify nothing has changed since
publication, hash the snapshot and compare:
curl -s https://cleartracedata.com/static/data/scorecard-edition-3.json | shasum -a 256
0c36db6ecc481a1d60ed4170d79b3f4d7f4c6b385ba399742c05617de15f118c
The snapshot is committed to a public git history at
publication time, which independently timestamps it.
Methodology & scope
Full methodology: cleartracedata.com/methodology and the
open-dataset README. In brief: execution slippage is the
median per-fill gap vs a 1-minute VWAP oracle over all on-chain trades; revert rate counts
failed routing transactions from raw on-chain data; quote accuracy is a forward-captured
fork-simulation sample (quoted vs realized). Aggregators we don't yet quote-sample appear
with on-chain metrics only.
Preliminary cells are shown but never rated. Quote-sample windows differ by venue (each
venue is rated once it has enough samples of its own), so the quote column is not a single
shared window the way the on-chain columns are. Scope: Ethereum for this edition; the live
leaderboard is the free preview layer that accrues between
editions.
Post-publication note (2026-09-25). The sentence above misnames the slippage baseline.
It says execution slippage is measured against “a 1-minute VWAP oracle.” It is not. The query
values both legs of each fill against Dune’s prices.usd at the minute of the fill, and
Dune documents that table as a view of its Coinpaprika feed: exchange-aggregate market data at
5-minute intervals, interpolated to the minute. No volume weighting is applied anywhere in the query.
Every figure on this page was computed exactly as it says elsewhere; only the name of the reference
price was wrong. Logged in our corrections log and registered in our metric-findings registry as
exec-slippage-baseline-is-not-a-vwap-2026-09-25.
Cite or commission
Cite this edition as: ClearTrace DEX Aggregator Execution Scorecard, Edition 3
(September 2026), sha256:0c36db6ecc48…. Link: https://cleartracedata.com/scorecard/edition-3. For a private cut of your own routing
versus peers (per-pair, per-size, with the failure modes), or a subscription to future rated
editions for a best-execution evidence file,
book a working session.