pyiceberg
Here are 20 public repositories matching this topic...
Lakevision is a tool which provides insights into your Apache Iceberg based Data Lakehouse.
-
Updated
Aug 6, 2026 - Python
Sample code to collect Apache Iceberg metrics for table monitoring
-
Updated
Aug 18, 2024 - Python
A poc open framework to manage data ingestion into apache iceberg tables
-
Updated
Sep 11, 2024 - Python
Official companion repository for the "Unleash the power of Apache Iceberg on AWS" Book. Includes code samples, deployment guides, ETL examples, and best practices for implementing scalable Apache Iceberg lakehouses on AWS.
-
Updated
Nov 26, 2025 - Jupyter Notebook
Tansu schema-backed topics, instantly accessible as Apache Iceberg tables with pyiceberg
-
Updated
May 31, 2025 - Just
-
Updated
Jan 23, 2026 - Go
Model Context Protocol (MCP) server for Apache Polaris. Enables AI agents and LLMs to interact with Polaris Catalog Management, audit data governance and access control (RBAC) policies, and perform Iceberg table inspections using PyIceberg.
-
Updated
Jul 5, 2026 - Python
Apache Iceberg with PyIceberg (no Spark/JVM): catalog setup, time travel, CDC, schema evolution, upserts
-
Updated
May 1, 2026 - Python
Apache Iceberg feature store and data lakehouse on Backblaze B2 — versioned feature tables, schema evolution, snapshot rollback, and DuckDB time-travel queries via PyIceberg, with B2 as the warehouse (no data warehouse to run)
-
Updated
Aug 18, 2026 - TypeScript
Apache Polaris + PyIceberg で Iceberg REST Catalog を動かして確かめた実験記録(2026年7月時点)
-
Updated
Aug 1, 2026 - Shell
provenance and immutability checks for iceberg lakehouses
-
Updated
May 21, 2026 - Python
Locally reproducible Apache Iceberg lakehouse with read-only FastAPI/DuckDB serving, near-real-time micro-batch ingestion, and a read-only NL agent over the API. Containerized with CI/CD, single-host go-live (Caddy auto-HTTPS), and OpenTelemetry→Datadog observability. Live synthetic demo: https://demo.agentic-data-platform.me/docs
-
Updated
Jul 11, 2026 - Python
Postgres CDC → Kafka → Apache Iceberg with exactly-once, out-of-order merge (source-LSN high-water-mark) and a source-to-lakehouse reconciler that quantifies drift in dollars. Ships a Spark Structured Streaming job, a docker-compose CDC stack, and Terraform.
-
Updated
Jul 8, 2026 - Python
A low-cost online Apache Iceberg lab using Cloudflare R2 Data Catalog
-
Updated
Aug 7, 2026 - Python
Iceberg table maintenance -- compaction, Z-order, partition evolution, snapshot expiry and orphan-file removal -- without needing Trino or Spark.
-
Updated
Aug 21, 2026 - Python
Apache Iceberg lakehouse: Bronze/Silver/Gold layers over a Yodlee-format transaction feed and equity prices, with schema evolution, cross-version queries, purged-CV return forecasting, a LangGraph ops agent, and scheduled data-quality and arrival monitors.
-
Updated
Aug 10, 2026 - Python
Provides a pyiceberg.io.FileIO implementation that uses hdfs-native client.
-
Updated
Aug 27, 2025 - Python
Data Lakehouse implementations showcasing two approaches: one built with Open Source technologies and another with Databricks
-
Updated
Sep 13, 2025 - Jupyter Notebook
Improve this page
Add a description, image, and links to the pyiceberg topic page so that developers can more easily learn about it.
Add this topic to your repo
To associate your repository with the pyiceberg topic, visit your repo's landing page and select "manage topics."