Skip to content
View muzammil-13's full-sized avatar
πŸ’­
a perpetual beginner with many half-built repos...
πŸ’­
a perpetual beginner with many half-built repos...

Block or report muzammil-13

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
muzammil-13/README.md

Hi there πŸ‘‹ I'm Muzammil

Data Engineer in Progress | DataOps Practitioner | Python β€’ SQL β€’ Snowflake | IBM Enterprise

muzammil-13

I'm a DataOps Analyst transitioning into Data Engineering, working in an IBM enterprise environment on automation, data validation, reporting workflows, and operational data quality.

My work sits at the intersection of Python, SQL, data validation, ETL automation, and cloud data platforms, where I solve data and automation problems at scale. I use AI-assisted development to accelerate exploration and implementation while deliberately building my own engineering fundamentals.

I began my career with a software engineering background, spent time as a developer intern at IBM, and gradually transitioned into enterprise DataOps, discovering a passion for building reliable data systems and developer tooling.

My goal: Build data infrastructure and automation that removes operational friction and enables better decision-making.


πŸš€ Current Focus

  • Enterprise DataOps: Building data validation, reconciliation, and reporting automation in an IBM environment
  • Data Engineering: Learning distributed data systems, orchestration frameworks, and cloud data platforms
  • Cloud Data Platforms: Working with Snowflake, Teradata, and cloud-native data architectures
  • AI-assisted Engineering: Using AI to accelerate development, debugging, and documentation while staying grounded in fundamentals
  • Batch Orchestration & Monitoring: Building automation and monitoring solutions for data pipelines and operational workflows

πŸ’» What I Build

Data Validation & Quality Automation

  • Automated data reconciliation and quality checks
  • ETL pipelines with error handling and logging
  • Reporting automation and distribution workflows
  • Batch job monitoring and operational alerting

Operational Automation

  • Python automation for repetitive data tasks
  • Data pipeline orchestration and monitoring
  • Internal tooling for enterprise data teams and analysts
  • Excel and SQL-based reporting automation

Data Engineering Foundations

  • Exploring distributed data processing (Apache Spark)
  • Learning workflow orchestration (Apache Airflow)
  • Building scalable data pipelines
  • Data-driven debugging and monitoring

πŸ“Œ Featured Projects

βš™οΈ Batch Monitor Agent

Internal enterprise automation project for monitoring batch-processing workflows, analyzing execution output, detecting operational conditions, and reducing manual investigation through automated detection and alerting.

As the technical lead and product owner for the end-to-end build, I'm using an AI-assisted SDLC approach to plan, design, implement, test, document, and iterate. The project demonstrates:

  • Python automation and batch job monitoring
  • Log and output analysis
  • Rule-based status detection and alerting
  • Operational workflow automation
  • Product-oriented engineering and technical ownership
  • AI-assisted design, implementation, and documentation

What it shows: automation thinking, operational systems, product ownership, AI-assisted engineering

Tech Stack: Python β€’ Batch Processing β€’ Automation β€’ Monitoring β€’ AI-assisted SDLC

(Implementation details are kept high-level because this is an internal enterprise project.)


πŸ” Operational Data Validator

A configurable validation engine for operational data workflows that turns repetitive manual data checks into repeatable, auditable validation workflows using reusable business-rule presets.

This project demonstrates:

  • Rule-based validation for data equality and tolerance checks
  • Preset-driven validation architecture
  • CSV and Excel input handling
  • PASS / WARN / FAIL validation outcomes
  • Reusable logic for data quality checks and reconciliation workflows
  • Python testing with pytest

What it shows: data validation fundamentals, modular architecture, data quality engineering, automation-first thinking

Tech Stack: Python β€’ Pandas β€’ PyYAML β€’ pytest β€’ Excel/CSV Validation


πŸ“Š Data Reconciliation Copilot

A local-first B2B MVP for comparing two CSV/XLSX datasets, finding deterministic discrepancies, using AI to explain the results, and exporting reconciliation reports.

The system combines:

  • Deterministic reconciliation logic using Pandas
  • Missing record, duplicate, mismatch, and schema issue detection
  • AI root-cause analysis and summary generation
  • React + FastAPI full-stack architecture
  • Supabase-ready auth and SaaS-ready structure

What it shows: full-stack data engineering, reconciliation workflows, AI-assisted root-cause analysis, deployment thinking

Tech Stack: React β€’ TypeScript β€’ FastAPI β€’ Pandas β€’ Supabase β€’ OpenAI


πŸ₯ Healthcare Claims Reporting Pipeline

A production-inspired Python ETL pipeline that automates end-to-end claims reporting from raw operational data to validated Excel reports and automated email delivery.

This project demonstrates core data engineering concepts:

  • Data validation and transformation
  • Error handling and logging
  • Scheduled automation
  • Operational report generation

What it shows: ETL design, data quality checks, reporting automation, Python best practices

Tech Stack: Python β€’ Pandas β€’ ETL β€’ Data Validation β€’ Automation β€’ Reporting


βš™οΈ Healthcare Claims Simulator

Simulation environment that mimics a real-world healthcare claims adjudication workflow. Used for understanding complex data workflows, experimentation, and learning data engineering patterns.

What it shows: system design thinking, workflow orchestration, data flow modeling, KPI monitoring

Tech Stack: Python β€’ FastAPI β€’ SQLite β€’ Pandas β€’ Streamlit β€’ Plotly


🧠 ExamInsights

RAG-powered platform that helps students analyze and extract insights from previous exam papers using Pathway and LLMs.

Explores LLM-assisted data processing and semantic search for document understanding.

What it shows: LLM integration, semantic search, containerized applications

Tech Stack: Python β€’ Pathway β€’ Streamlit β€’ Docker β€’ LLM


πŸ” FileShazam (Private)

Semantic desktop search engine for finding files by meaning instead of filename.

What it shows: embeddings, semantic search, file system automation

Tech Stack: Python β€’ Embeddings β€’ Semantic Search


πŸ›  Technical Stack

Core Languages

Python β€” automation, data processing, ETL, scripting
SQL β€” data modeling, analysis, query optimization, validation
Bash β€” scripting and operational automation

Data Engineering & DataOps

  • ETL/ELT: Data pipeline design, transformation, validation, error handling
  • Data Quality: Reconciliation, schema validation, automated data checks
  • Data Platforms: Snowflake, Teradata, PostgreSQL
  • Python: Pandas, data manipulation, automation libraries
  • Batch Processing: Scheduling, monitoring, orchestration fundamentals

Cloud & Platform Tools

  • Snowflake β€” cloud data warehousing, query optimization
  • Teradata β€” enterprise data warehouse, SQL fundamentals
  • PostgreSQL β€” relational database design and optimization
  • Docker β€” containerization and reproducible environments
  • Azure β€” cloud services and enterprise infrastructure
  • Git/GitHub β€” version control, collaboration

AI & Development

  • AI-assisted development β€” code generation, debugging, documentation with AI
  • LLM integration β€” RAG, semantic search, document processing
  • Streamlit β€” rapid internal tool development

Currently Learning

  • Apache Airflow β€” workflow orchestration and scheduling
  • Apache Spark β€” distributed data processing
  • Advanced System Design β€” scalable data systems
  • Cloud Data Engineering β€” managed data pipelines and platforms

πŸ† Certifications

IBM Certified Data & AI (CCDVF) β€” Completed, Q4 2026


🌱 Engineering Philosophy

AI-assisted, fundamentals-driven.

I use AI (LLMs, code generation, documentation assistance) to accelerate exploration, implementation, and debugging. But AI is a multiplier for fundamentals, not a substitute for them.

My approach:

  • Use AI to rapidly prototype and explore
  • Deliberately understand the underlying systems and design decisions
  • Build with modularity, error handling, and scalability in mind
  • Learn in public and document what I discover
  • Prioritize shipping useful internal tools over experimental prototypes

Core principles:

  • Build useful software
  • Automate repetitive work
  • Think in systems
  • Improve through consistency
  • Ship incrementally, learn continuously

πŸ’Ό Experience

DataOps Analyst (IBM Enterprise Environment)

Working on enterprise data operations, automation, and reporting workflows. Responsibilities include:

  • Python automation and data validation scripting
  • SQL development and data quality checks
  • Data pipeline monitoring and operational analytics
  • ETL process design and optimization
  • Batch job scheduling and orchestration
  • Reporting automation and delivery
  • Exposure to Snowflake and Teradata migrations

Previously completed a developer internship at IBM, where I transitioned into DataOps and discovered a passion for solving data and automation problems at scale.

Learning trajectory: Software Engineering β†’ IBM Intern β†’ Enterprise DataOps β†’ Data Engineering (current focus)


🌐 Connect


πŸ“Š GitHub Activity


Building data systems that remove friction, enable automation, and create lasting operational impact.

Popular repositories Loading

  1. Django_projects-inmakes Django_projects-inmakes Public

    This repository showcases real-world implementations across different domains using Django framework.

    CSS 4

  2. muzammil-13 muzammil-13 Public

    My GitHub Readme

    2

  3. data_analysis-inmakes data_analysis-inmakes Public

    A data-driven project that leverages machine learning to predict Bitcoin price trends. Using historical Bitcoin data, this analysis provides 30-day price forecasts through advanced statistical mode…

    Jupyter Notebook 2

  4. muzammil-13.github.io muzammil-13.github.io Public

    My Personal portfolio hosted through Github pages.

    JavaScript 2

  5. Calculator-sample Calculator-sample Public

    A Sample Reponsive calculator using Javascript.

    CSS 2

  6. ExamInsights_app ExamInsights_app Public

    ExamInsights helps students and educators analyze past exam papers with a Pathway-powered RAG backend and a Streamlit interface.

    Python 2 1