Glue scripts for converting AWS Service Logs for use in Athena
-
Updated
Feb 1, 2024 - Python
8000
Glue scripts for converting AWS Service Logs for use in Athena
The Sensitive Data Protection on AWS solution allows enterprise customers to create data catalogs, discover, protect, and visualize sensitive data across multiple AWS accounts. The solution eliminates the need for manual tagging to track sensitive data such as Personal Identifiable Information (PII) and classified information.
Build and deploy a serverless data pipeline on AWS with no effort.
Extract, transform, and load data for analytic processing using AWS Glue
This is a data pipeline built with the purpose of serving a business team.
Terraform configuration that creates several AWS services, uploads data in S3 and starts the Glue Crawler and Glue Job.
Terraform module which creates Glue Job resources on AWS.
First experimentation into the world of AWS Serverless Data Engineering
Pipeline ETL na AWS
A cloud-based ETL pipeline on AWS for automating airline flight data ingestion, transformation, and storage using S3, Glue, Redshift, EventBridge, Step Functions, and SNS.
This project demonstrates an end-to-end pipeline for ingesting CloudWatch Logs into Apache Iceberg tables and building materialized views using AWS Glue.
🌾 AWS Serverless ELT: Agricultural & weather data pipeline with automated quality checks and Athena/BI integration.
AWS Glue is a fully managed, serverless data integration service that makes it easy to discover, prepare, and combine data for analytics, machine learning, and application development. It automates the ETL (Extract, Transform, Load) pipeline process, allowing users to focus on building workflows rather than managing infrastructure.
This project outlines the final project requirements for Information Architectures, focusing on group assignments, scoring criteria, topic selection, core requirements, and project components such as design, development, visualization, and executive presentation.
ELT (Extract, Load, Transform) pipeline that fetches stock data from Yahoo Finance, stores it in an S3 bucket, and then loads it into an Redshift Serverless table
Data Engineering project using data streaming produced by python applications, ETL process and availability for ad-hoc SQL queries in the AWS cloud
Exports CloudWatch Log Groups to S3 in a structured, date-partitioned format using an AWS Glue Python Shell job.
Add a description, image, and links to the glue-job topic page so that developers can more easily learn about it.
To associate your repository with the glue-job topic, visit your repo's landing page and select "manage topics."