Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.
-
Updated
Dec 21, 2024 - Rust
FFFF
Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.
The simplest way to scale Python.
Data pipelines from re-usable components
The open-source Useful SDK. One python decorator in the Useful library allows for full observability of Python functions within an ETL.
Data Cleaning for Pyspark
A project structure for doing and sharing data engineer work.
Lien de l'application
Master the AWS Data Stack! 🚀 This repository features 15+ Industrial Data Engineering Projects covering Serverless ETL, Real-Time Streaming, & Data Warehousing. Hands-on labs for S3, Lambda, Spark, Airflow, Snowflake, Redshift, Kinesis, & Glue. Includes production-grade CICD pipelines. A complete roadmap to becoming a top Data Professional.
End To End MLOPS Project With ETL Pipelines- Building Network Security System
A smart parking decision-support platform for drivers unfamiliar with Melbourne CBD. Built for FIT5120 Industry Experience Studio (S1 2026) by Team FlaminGO.
A Data Pipeline and Data House For Euro Flight Data.
An end-to-end cloud data pipeline tracking 400+ AI models from OpenRouter. Features automated Python ETL transformations, robust relational validations, a Supabase PostgreSQL cloud warehouse, and a premium glassmorphic live analytics dashboard detailing model context and price benchmarking.
This repository contains my first end-to-end Data Engineering project, built using Microsoft Azure Cloud and Azure Databricks with PySpark.
Python ETL pipeline normalizing Ahmedabad transit APIs (BRTS, AMTS, Metro) into unified JSON/GTFS datasets for GraphHopper and OpenTripPlanner
DataSift auto applies a data pre-processing pipeline to Data Science Projects.
End-to-end Azure Data Factory project transforming raw sales data into customer-level insights using pivot transformation and storing results in Blob Storage.
Complete portfolio of data engineering projects from Udacity's Data Engineering with AWS Nanodegree.
A robust ETL pipeline using MySQL for data transformation and Python for programmatic visualization of the 2018 NYC Squirrel Census (3,000+ records).
Build ETL piplines on AirFlow to load data from BigQuery and store it in MySQL
🗄️ IBM Relational Database Administrator with GenAI Certificate Portfolio – A comprehensive collection of projects, labs, and assignments showcasing expertise in relational database administration, 🏘️data warehousing, 🔁ETL pipelines, and 🤖Generative AI integration for modern database management.
Add a description, image, and links to the etl-pipelines topic page so that developers can more easily learn about it.
To associate your repository with the etl-pipelines topic, visit your repo's landing page and select "manage topics."