Cloud Computing and Data Centers

Amazon Aurora PostgreSQL Enables Direct Queries from Iceberg and Parquet Within the Data Lake

AWS has added to Amazon Aurora PostgreSQL the ability to combine live operational data with Apache Iceberg and Parquet data stored in Amazon S3 in a single PostgreSQL query, without ETL pipelines to move the data. The capability relies on embedded DuckDB and supports the AWS Glue Data Catalog and Iceberg catalogs compatible with IRC.

2026-09-30
4 min read
18 views
certi.news Editorial Team
Amazon Aurora PostgreSQL Enables Direct Queries from Iceberg and Parquet Within the Data Lake

AWS announced the addition of the ability to query Apache Iceberg and Apache Parquet data stored in the data lake alongside operational data within Amazon Aurora PostgreSQL. The new capability allows applications to combine recent and historical records in a single query using PostgreSQL syntax and existing tools, without copying the data into the database through ETL pipelines.

What has changed in practice?

The functionality relies on the DuckDB engine embedded directly within Aurora PostgreSQL. As a result, a single query can read operational data, including uncommitted writes, and read Iceberg tables or Parquet files from Amazon S3. It also supports the AWS Glue Data Catalog, Amazon S3 and S3 Tables, in addition to Iceberg REST Catalog-compatible catalogs through the federation mechanism in Glue.

The user needs to create an Aurora PostgreSQL cluster, attach an IAM role that includes the AuroraAnalytics feature, and then enable the aurora_analytics extension. After that, foreign tables can be created to point to external data, or IMPORT FOREIGN SCHEMA can be used to create multiple tables automatically while inferring the data schema from metadata.

Why does this news matter?

Applications that combine recent transactions in Aurora with historical records in S3 have typically needed to copy data through reverse ETL, adding synchronization processes, operating costs, and engineering complexity. Now, this integration can be performed during the query itself. This benefits dashboards that need historical context, applications that link a transaction to a previous record, and systems that handle data distributed between an operational database and a data lake.

This point is especially important for applications that use artificial intelligence agents, because it is difficult to predict in advance every dataset an agent may need and copy it into an operational database. Direct access makes it possible to expand the available data without creating a copy for every use case.

Performance and access options

Aurora applies optimizations such as pushing filters to the data source and reducing the columns read, while frequently used data is temporarily stored inside the Aurora instance. The behavior of each query can be examined through the aurora_analytics_stat_statements() function, which displays metrics including the number of rows scanned, the volume of data read from S3, and cache utilization.

Queries are available on the various Aurora instances within the cluster, including the writer and read replicas, allowing analytical scanning operations to be assigned away from the operational workload. For cases requiring response times on the order of fractions of a millisecond, data can be materialized in a native Aurora table using CREATE TABLE AS SELECT, INSERT INTO ... SELECT, or MERGE INTO. Writes resulting from materialization are executed on the writer instance.

Availability and cost

The capability supports Aurora PostgreSQL 17 starting with 17.11, and 18 starting with 18.6. It is available in all AWS commercial Regions and AWS GovCloud (US) Regions, with no additional fee for the feature itself. Costs remain associated with the additional Aurora compute consumed by the queries, in addition to Amazon S3 request charges for reading data lake files.

certi.news reading: The fundamental change is not the addition of a new storage format, but rather reducing the distance between the operational database and the data lake. However, this does not eliminate workload-design considerations: direct querying is suitable for flexible data access, while materialization remains necessary when extremely low response times are a priority. Aurora compute consumption and S3 requests will also remain factors that must be measured according to the query pattern and data volume.

News source
AWS News Blog
Open original source ↗
c
Author

certi.news Editorial Team

In the same category

You may also like

View all news