Polars has released version 2.0 of its DataFrame engine, bringing broader SQL support and changing the default execution behavior of lazy queries. The project says the release is intended to serve more SQL workloads while making large analytical jobs more resilient when memory is constrained.

The most consequential compatibility change is that collecting a LazyFrame now uses the streaming engine by default. Streaming can reduce memory use and improve performance, but some operations—including joins, grouping and unpivoting—do not preserve row order unless users request it. Applications that depend on observable order must set `maintain_order=True`, one reason the project treated the update as a major-version release rather than an ordinary feature upgrade.

Out-of-core execution is also enabled by default for supported operations. Polars begins spilling work to disk at around 80% of available memory and uses a default disk budget of 64GB. Sorting, window functions and many expressions can currently use the mechanism. Out-of-core joins and group operations remain on the roadmap, so the protection does not yet cover every memory-intensive query.

SQL is now described by the project as a first-class interface. Recent work expanded statement coverage and added engine optimizations including join reordering, stronger common-subplan elimination, dynamic predicates and Bloom filters. Polars published benchmark results based on data derived from TPC-H and TPC-DS, comparing its SQL engine with DuckDB and DataFusion on 16-vCPU and 192-vCPU Amazon instances. The project reported leading most of its selected comparisons, but emphasized that the tests do not comply with official TPC benchmark rules. It also published the benchmark repository so others can reproduce or challenge the results.

The release adds direct support for Arrow’s MapType through a Polars Map data type. Previously, Arrow map values were represented as lists of key-value structures. Dedicated support is intended to enable dictionary-style expressions such as key lookup and iteration over values.

Polars is also leaning into stricter validation. The project argues that mismatched data should produce explicit errors instead of silent coercion, especially as AI coding tools generate more queries. Developers and automated agents can call `collect_schema()` to resolve types and catch some structural problems before a query materializes data, although data-dependent failures can still occur during execution.

The 2.0 label therefore marks both new capabilities and behavioral changes that require attention during migration. Users gain streaming by default, spill-to-disk support and a wider SQL surface, but should review assumptions about ordering and unsupported out-of-core operations. The project has published a migration guide and says its next priorities include broader disk spilling, better scaling on high-core-count machines and continued work on distributed execution.