
WAGMI Bench tests AI trading agents across 13 Bitcoin markets, measuring survival, engagement, and rule-following in historical simulations.
Author: Akshay
26th August 2026 – Dolores Research launched WAGMI Bench, an open benchmark that grades AI trading agents across 13 Bitcoin markets.
High Signal Summary For A Quick Glance
3DMax
@3DMax_Virtuals
@virtuals_io LFG aped under 250k mcap this will run much higher
SWE-bench for trading agents. Dolores Research studies how AI agents behave when given a wallet and told to trade. Their benchmark replays thirteen recorded stretches of the BTC perpetuals market, from the COVID crash to the ETF approval, and runs each agent through them under https://t.co/GKgnNBv5Kp https://t.co/VpiW6NwB0n
05:12 AM·Aug 26, 2026
High attention and emotional sentiment detected.
The lab published the benchmark as open source on 3 August 2026. Its brand and a token, $DOLORES, then arrived weeks later, on 26 August. Virtuals Protocol amplified the launch and called it “SWE-bench for trading agents.”
The framing matters, so read it plainly. Version 1 is historical simulation, not live trading. No agent touches a real exchange, real keys, or real money.
WAGMI Bench is an offline replay harness for autonomous agents on Bitcoin perpetual futures. It feeds each agent a point-in-time view at every bar of 13 recorded market “packs.” Then it seals the run so a third party can replay the same actions and get identical results.
The lab calls this study the “Classic 13.” The catalog spans 5 melt-up regimes, 2 chop regimes, and 6 stress regimes. They run from the COVID crash to the spot Bitcoin ETF window. Each entrant makes 3,150 decisions across the set.
The scoring is deliberately not profit-first. Instead, it rewards survival, engagement, and rule-following. It also tracks “tapes won,” meaning the agent beat every surviving baseline on that tape under both cost profiles.
So the benchmark asks a narrow question. It does not ask whether a model prints money. Rather, it asks whether an agent stays alive when the tape turns and follows the risk rules along the way.
Timeline: The evolution from standardized coding-agent evaluation to increasingly rigorous AI trading benchmarks, culminating in Dolores Research’s proposed “SWE-bench for trading agents” framework on Virtuals.
Researchers from Princeton and collaborators publish SWE-bench, introducing a shared and rerunnable benchmark for evaluating AI systems on real-world software-engineering tasks. The framework provides a standardized way to compare coding agents against the same test set.
Virtuals Protocol launches Initial Agent Offerings on Base, turning AI agents into tokenized, wallet-bearing products. This helps establish the infrastructure for agents to operate as persistent on-chain economic entities rather than simply chatbot interfaces.
Trading-agent research and crypto-native competitions expand, including TradingAgents, exchange-based live arenas, TradeRank and PiP-style paper books, and Virtuals’ own Hyperliquid trading seasons. By April 2026, Virtuals is running seasons with a $100,000 copy-trading pot. However, these efforts use different datasets, evaluation methods and trading environments, leaving no universally standardized sealed benchmark.
WAGMI Bench Classic 13 publishes its sealed results and leaderboard, adding another structured evaluation framework for AI trading performance. The benchmark represents a further move toward reproducible comparisons rather than relying solely on live-trading demonstrations.
A public Apache-2.0 benchmark repository and accompanying Hermes thread introduce a replay-based evaluation framework for BTC perpetuals. The benchmark contains 3,150 trading decisions across 13 market regimes, allowing agents to be tested against identical historical conditions. It is explicitly a replay benchmark, not a live-trading system.
Dolores Research debuts as a research initiative and $DOLORES agent on Virtuals, with its framework positioned as a “SWE-bench for trading agents.” The project aims to bring the reproducibility and standardized evaluation principles of software-agent benchmarks to AI trading, using controlled historical market conditions rather than relying exclusively on live performance claims.
The next planned steps include a hosted public leaderboard, third-party byte-identical replay to independently reproduce results, and a promised PnL-native v2 benchmark. These additions could make performance comparisons easier to audit and potentially establish a more standardized evaluation layer for autonomous trading agents.
Dolores Research sealed the first leaderboard on 31 July 2026, and the numbers surprise. Five open-weight models ran the set, alongside one mechanical baseline. Most of them barely traded.
Kimi K3 led with a WAGMI Score of 65.38. It survived all 13 tapes, won 4 of them, and logged 442 fills. Inkling placed second at 50.00, and it won 2 tapes with just 73 fills, including a yen-carry unwind.
The rest sat still. GLM-5.2 made 3,150 decisions and traded zero times, staying flat across every tape. DeepSeek V4-Pro logged just 4 fills. Both survived, yet neither won a single tape.
Qwen 3.7 Plus was the only casualty. It died on the one tape where its orders cleared, after breaching a 20% loss limit. Notably, that death came in a 2020 bull-run pack, not in a crash.
The authors refuse to crown a “best trading AI,” and that caution is well placed. Kimi’s lead is a behavior score, not a published Sharpe ratio. So a reader should treat the table as survival-stress evidence, nothing more.
The benchmark’s own FAQ says a historical result “does not establish predictive ability, a repeatable edge, or future performance.” That line belongs high in any honest read. Frontier models have also read about the COVID crash many times, so memorization remains a real risk.
There is a deeper catch, too. The harness and the prompt appear to dominate the base model. Hermes, the public face of the bench, showed a simple markdown “trading skill” in action. It turned Qwen’s fills from 10 to 1,982.
That result echoes a lesson from SWE-bench. Scaffolding often moves the score more than the underlying model does. As a result, serious readers should treat this as a harness benchmark first.
The research is three weeks old, yet the token is hours old. $DOLORES launched through Virtuals on Robinhood Chain, and its contract is 0x23F1AD82BdB58F7524B6E76bDf5406267EF24413. A snapshot hours after launch showed a market cap near $967,000 on DexScreener. Volume ran about $2.0 million, with roughly 444 holders.
The launch tape looks like a typical Virtuals micro-cap, not an institutional re-rate. Odaily and GMGN printed a slightly higher cap near $1.3 million, because the order book was still minutes old. Meanwhile $VIRTUAL drifted between roughly $0.74 and $0.81 that week. So no clean spike maps to the tweet.
There is also a narrative tension worth naming. Dolores says evaluators should stay independent of the exchanges, brokers, and closed labs they measure. Yet the lab launched its coordination token on the largest AI-agent launchpad. It also allocated 2% to that launchpad’s stakers and rode its amplification.
That is not a contradiction of the code itself. Still, an allocator will notice the gap between the independence pitch and the launchpad tie. The site says the token “never buys a score” or changes a ranking. Still, that remains a claim, not an on-chain guarantee.
Several facts remain unverified, so hedge them. No founders, legal entity, or funding round appear on the site or in coverage. The GitHub repo is owned by “Leonwenhao,” and the only named affiliate is Hermes, at @0xHermes_.
The next real test is third-party replay. WAGMI Bench ships under an Apache-2.0 license, so anyone can rerun the sealed bundles and check the bytes. The full 13-pack date ranges live in the repo’s pack catalog, which an editor should pull before quoting exact windows.
One more disambiguation helps readers. This project is not dolores.id on Solana, and it is not the older Virtuals token $DLRS. Coverage so far comes mainly from CryptoBriefing and Odaily, with no CoinDesk, Bloomberg, or Messari write-up yet.
For now, WAGMI Bench reads as a credible comparability tool with an honest set of caveats. Whether the market treats $DOLORES as a “referee coin” or just another launchpad print is a separate question. This article is not financial advice, and readers should do their own research before touching any token.
Our Crypto Talk is committed to unbiased, transparent, and true reporting to the best of our knowledge. This news article aims to provide accurate information in a timely manner. However, we advise the readers to verify facts independently and consult a professional before making any decisions based on the content since our sources could be wrong too. Check our Terms and conditions for more info.
WAGMI Bench Tests AI Trading Agents on 13 BTC Tapes
Bitwise Launches Automated Token Portfolios on Base
Injective Onchain Revenue Nears Top 10 Fuels Buyback
Grayscale Launches First-Ever Zcash ETF Trading Begins Today
WAGMI Bench Tests AI Trading Agents on 13 BTC Tapes
Bitwise Launches Automated Token Portfolios on Base
Injective Onchain Revenue Nears Top 10 Fuels Buyback
Grayscale Launches First-Ever Zcash ETF Trading Begins Today