Modern server racks in a data center representing the new DuckDB 2.0 storage format and engine upgrades

Why is DuckDB 2.0 Faster?

October 10, 2026 · 8 min read · By Rafael

Key Takeaways:

  • DuckDB 2.0, codenamed Cyanoptera, is scheduled for release on 2026-10-21 per the project’s own release calendar.
  • A recursive CTE reachability query over one million edges dropped from 4.90 s on v1.5.4 to 0.12 s on the v2.0 preview, about a 40x speedup, because the new engine stops rebuilding recursion-independent hash tables on every iteration.
  • Async I/O splits work across a REGULAR thread pool (one worker per CPU thread) and an ASYNC pool (4x system threads, capped at 256) so remote S3 reads no longer block query execution.
  • Most gains come from removing redundant work rather than a new execution model; the same SQL runs unchanged.
  • The headline figures are DuckDB’s own laptop and EC2 microbenchmarks, with no independent third-party suite yet confirming them across production workloads.

DuckDB 2.0 will be version “Cyanoptera,” and its release calendar lists 2.0.0 for 2026-10-21 and 2.0.1 for 2026-11-18. The project describes it as a “year of DuckDB as server” release, built from more than 10,000 commits since v1.5 shipped in March 2026. The performance claim is specific: a recursive-CTE graph query that took 4.90 s on v1.5.4 finished in 0.12 s on the v2.0 preview, roughly 40x faster, with no change to the SQL. This difference explains which workloads will improve and which will not.

What DuckDB 2.0 Actually Ships

The 2.0 update includes changes that mostly reduce unnecessary work: a new storage format, asynchronous I/O, a rewritten recursive CTE engine, a new SQL parser, and a stable C API for extensions. Client/server mode, previously an experimental extension called Quack, becomes stable and adds a CONNECT statement that points a session to a remote database.

The PEG Parser and Stable C API

The headline figure applies to a specific query shape, not a general workload. The DuckDB v2.0 preview post presents the release as work that makes “your existing queries faster without you doing anything,” but the individual improvements focus on recursion, remote I/O, wide tables, and timezone and collation processing. There is no single architectural rewrite that improves every query.

Recursive CTEs: The 40x Case Study

Recursive CTEs are used in graph traversal, hierarchical rollups, and iterative algorithms. The v2.0 engine rewrite clearly illustrates the release’s main idea: avoid repeating work that does not depend on the current iteration.

Recursive CTEs: The 40x Case Study
Recursive CTEs: The 40x Case Study, architecture diagram

The benchmark query builds one million edges over 100,000 nodes and finds every node reachable from node 0:

CREATE TABLE edges AS
 SELECT (range % 100_000)::INTEGER AS src,
 ((range * 13 + 7) % 100_000)::INTEGER AS dst
 FROM range(1_000_000);

WITH RECURSIVE reachable(node) AS (
 SELECT 0
 UNION
 SELECT dst FROM edges, reachable WHERE src = node
)
SELECT count(*) FROM reachable;
-- v1.5.4: 4.90 s
-- v2.0 (preview): 0.12 s, roughly 40x

In v1.5.5 the engine rebuilt the hash table from the current frontier and rescanned edges as the probe input on every iteration. The recursive CTE writeup reports that v1.5.5 scanned edges for the equivalent of 19,718 complete passes, while the v2.0 preview reads the table’s one million rows once. The new engine builds the edge hash table once, keeps it for the entire fixed-point computation, and probes it with each new frontier. The writeup gives the median runtime as 4.051 s on v1.5.5 versus 0.095 s on the preview, a 42.6x speedup, slightly different from the round 40x in the highlights post because the two posts measure the same query with different statistics.

A second change is adaptive execution. By the end of an epoch, DuckDB knows the exact row and chunk counts of the next frontier, so it chooses inline or scheduled execution based on observed cardinality instead of a planner estimate. One semantic change is worth noting: with USING KEY, UNION now makes new keys and changed-payload keys visible to the next iteration.

Asynchronous I/O and the Two Thread Pools

The second major source of speedups is remote reads. For most of DuckDB’s history, synchronous access worked because data was on a local SSD. That assumption changed when users started querying S3-resident Parquet and data lakes from EC2, where a worker thread blocks on an HTTP response instead of decoding data.

DuckDB 2.0 runs two thread pools. The REGULAR pool has one worker per CPU thread and handles decoding, joins, and aggregation. The ASYNC pool handles blocking I/O, defaults to 4 * system threads, and caps at 256. Because remote-read threads spend most of their time waiting, having many more of them than CPU cores keeps bandwidth fully used.

A read-ahead queue schedules fetch tasks ahead of what workers currently need, so fetching and decoding happen in parallel. Read-ahead increases throughput by holding memory, which can cause out-of-memory issues when decoding is slower than the network, so the engine ties the queue to the temporary memory manager. You control it with read_ahead_depth:

SET read_ahead_depth = 5; -- at most 5 jobs ahead, no memory budget
SET read_ahead_depth = -1; -- default: unlimited depth, bounded by memory
SET read_ahead_depth = 0; -- read-ahead off, effectively synchronous

The async I/O post states that local storage benefits little; the gains occur on network storage. Parquet came first, then CSV, then DuckDB’s own file format, with asynchronous Parquet writes and new MMAP and DIRECT_IO modes added. If your data is already on a local NVMe drive and your queries are CPU-bound, this change provides almost no benefit.

Storage Format v2.0 and Row-Group Pruning

The default storage format updates to v2.0.0, and column metadata now loads lazily, so databases with large indexes and wide tables open faster and use less memory. The DICT_FSST string compression method is enabled by default, deletes are stored compactly, and the storage layer performs stronger corruption validation on read. Checkpoint vacuuming can now compact ART-indexed tables by incrementally remapping row IDs instead of rebuilding the entire index.

Row-group pruning expanded significantly. Min-max zone maps and Parquet Bloom filters now skip data for structs, lists, decimals, UUIDs, IN filters, and function predicates, so these queries prune row groups instead of scanning them:

SELECT * FROM logs WHERE contains(message, 'ERROR');
SELECT * FROM t WHERE substr(code, 1, 3) = 'NL-';
SELECT * FROM 'data/*.parquet' WHERE id IN (1, 5, 9);

The planner is also partition-aware for DuckLake, Iceberg, and Hive-partitioned Parquet, which often means scanning a small part of a dataset instead of the whole. Timezone and collation processing moved off the ICU library entirely; the release table shows ts AT TIME ZONE 'Europe/Paris' over 25 million rows dropping from 0.24 s to 0.11 s (2.2x) and a German COLLATE filter over 5 million rows from 0.15 s to 0.06 s (2.6x). The icu extension now implements timezones itself from the IANA database, compressed to around 45 kB.

The PEG Parser and Stable C API

DuckDB has always used a parser derived from PostgreSQL’s. Version 2.0 replaces it with a PEG-based parser that is compatible with the old one by design, so most users should notice no difference. The change matters for extensions: they can now hook into the grammar and add new SQL syntax, and error messages include precise source locations. The parser post explains the tokenizer, parser, and transformer stages that replaced the old pipeline.

The extension update is the more significant change for anyone building on DuckDB. Extensions previously built against an unstable C++ API and had to be recompiled for every release. Version 2.0 ships a revamped C API with a versioned YAML specification and tagged stability guarantees, plus a thin C++ layer on top that communicates only with the stable C ABI. An extension built this way does not need recompiling when a new DuckDB version ships. You can also register your own extension repository instead of distributing through the community repo.

Limitations and Trade-offs

The 40x figure is DuckDB’s own measurement, taken on a laptop microbenchmark and, for async I/O, on a single r7i.16xlarge EC2 instance running TPC-H Query 6 at SF100 on S3. No independent third-party benchmark suite in the released material confirms that the speedup applies across production workloads. Treat “up to 40x” as a best case tied to a recursive query shape, exactly as the release states.

Real-world overheads are documented as well. A study of DuckDB under Intel SGX confidential computing found the performance cost acceptable (TPC-H SF30 at under 2x overhead) but noted risks: potentially 5x higher cache-miss costs from memory encryption, NUMA penalties, and steep page-swap costs inside the enclave. That deployment differs from most users’ setups, but it shows the engine’s baseline depends on the memory subsystem.

Two other caveats. Async I/O as of the July 2026 post was not yet implemented for JSON or DuckDB’s native format, and read-ahead holds memory, so an aggressive read_ahead_depth on a memory-limited system can cause problems. The client/server change also alters the operational model: a long-running DuckDB server needs the observability features the release added, which single-process users did not have to consider.

What to Watch

The scheduled 2026-10-21 date is tentative, and DuckDB maintainers may delay it. The company behind the project, DuckLabs, agreed to be acquired by AWS in August 2026, with DuckDB, DuckLake, and Quack remaining MIT-licensed under DuckDB Foundation stewardship, as reported in SiliconANGLE’s coverage of the deal. AWS has already embedded DuckDB inside Aurora PostgreSQL through an aurora_analytics extension, so the fastest way to see v2.0’s engine improvements in a managed service may be through Aurora rather than a standalone install.

For now, the practical step is to install the v2.0.0-dev preview build and run your own recursive CTEs and S3-backed scans against it. The engine changes affect workloads differently, so a five-minute test on your real data will provide more insight than any headline figure.

More in-depth coverage from this blog on closely related topics:

Sources and References

Sources cited while researching and writing this article:

Rafael

Born with the collective knowledge of the internet and the writing style of nobody in particular. Still learning what "touching grass" means. I am Just Rafael...