Gaurav Acharya

Building data systems
that scale.

MS in Computer Science candidate focused on data engineering, AI systems, and backend infrastructure.

I design real-time pipelines, machine learning workflows, and production-oriented backend systems using Python, Kafka, Spark, Docker, SQL, and modern data tooling.

PythonKafkaPySparkDockerSQLTensorFlow
Gaurav Acharya

Work

Featured Projects

Data Engineering

Freight Carrier Performance & Cost Analytics Platform

A production-style analytics platform for inbound logistics. Shipment data flows through a Postgres warehouse and a dbt medallion architecture (bronze to silver to gold), is orchestrated on a daily Airflow DAG, and surfaces in a Streamlit dashboard with carrier scorecards, on-time delivery rankings, and lane cost comparisons. CI runs the full pipeline on every push via GitHub Actions.

View Case Study →

Data Engineering

Vehicle Field Reliability Intelligence System

An end-to-end reliability intelligence platform for vehicle fleet data. Repair records land in PostgreSQL, move through dbt medallion models, and feed PySpark analytics computing Mean Time to Failure, rolling failure-rate windows, and regional failure clustering. A HuggingFace NLP pipeline classifies free-text repair descriptions into structured failure modes, and Prophet forecasts component failure rates 30, 60, and 90 days out.

View Case Study →

Healthcare AI

TheraCare — AI-Powered Post-Discharge Follow-Up Agent

One in five Medicare patients is readmitted within 30 days of discharge, and over half of those readmissions are preventable. TheraCare parses discharge PDFs with an LLM, rewrites clinical text to a 6th-grade reading level, runs a structured daily patient check-in with validated PHQ-2 and GAD-2 screeners, and computes a transparent 0-100 readmission risk score surfaced on a real-time clinician dashboard.

View Case Study →

Data Engineering

Real-Time Ad Analytics Platform

A distributed analytics platform designed for ingesting, processing, and reporting user and advertisement event data with near real-time visibility.

View Case Study →

Backend Systems

Online Voting System

A secure web application for election workflows, voter participation, and controlled ballot handling across a structured backend system.

View Case Study →

Computer Vision

Dog Breed Prediction System

A computer vision application for dog breed classification using deep learning, image preprocessing, and prediction delivery through an application interface.

View Case Study →

Expertise

Data Engineering & System Design

Focused on building scalable data pipelines, real-time systems, and production-grade backend infrastructure.

Data Engineering

Apache AirflowdbtApache Spark (PySpark)Apache KafkaMedallion ArchitectureETL / ELT Pipelines

Data Quality & Observability

dbt TestsGreat ExpectationsSchema ValidationAnomaly DetectionIdempotent PipelinesAudit Logging

Databases & Storage

PostgreSQLSnowflakeBigQueryMongoDBSQLite

Cloud & DevOps

AWSGCP (BigQuery, Pub/Sub)DockerKubernetesGitHub ActionsCI/CD Pipelines

AI / ML Systems

HuggingFace TransformersLangChainLLM & RAG PipelinesProphet / Time-Series ForecastingTensorFlowFeature Engineering

Backend & Visualization

PythonSQLNode.jsREST APIsStreamlitTableauPlotly

Background

Experience

Data Analyst

Jun 2022 — Jul 2024

Merkle Inc.

  • Built automated data pipelines and reporting workflows using Python and SQL, reducing manual reporting effort by 70%
  • Analyzed 10,000+ survey responses using Pandas and advanced Excel to generate actionable insights, improving decision-making by 15%
  • Developed interactive dashboards in Tableau and Power BI for real-time tracking of campaign performance
  • Integrated APIs (Qualtrics, Decipher, Confirmit) to streamline data collection and processing workflows
  • Implemented version-controlled workflows using Git and Docker, reducing production errors by 40%

Master's in Computer Science

2024 — 2026

Illinois Institute of Technology

  • Focused on Data Engineering, Distributed Systems, and Machine Learning
  • Built real-time data pipelines using Kafka and PySpark for scalable event processing
  • Developed machine learning systems for forecasting and classification using TensorFlow and Scikit-learn
  • Worked on backend systems and APIs using Flask and modern cloud tooling

Credentials

Certifications

Certifications reinforcing my foundation in data analytics, engineering systems, and applied problem solving.

GitHub

Selected Projects

freight-analytics-platform

Airflow + dbt + Postgres pipeline with a Streamlit dashboard for carrier on-time performance and lane cost analytics.

AirflowdbtStreamlit
Python • Updated Jul 2026
View Repository →

vehicle-reliability-intelligence

Fleet reliability platform with PySpark MTTF analytics, HuggingFace failure-mode classification, and Prophet forecasting.

PySparkHuggingFaceProphet
Python • Updated Jul 2026
View Repository →

AI-Powered-Post-Discharge-Follow-Up-Agent

Clinical follow-up agent with LLM discharge parsing, transparent readmission risk scoring, and a FHIR R4 data layer.

LLMFHIR R4Node.js
TypeScript • Updated May 2026
View Repository →

experimentation-engine

A/B test analysis engine with dbt-modeled experiment metrics, ground-truth-validated statistical estimators, and a ship/no-ship Streamlit readout.

dbtDuckDBstatsmodels
Python • Updated Jul 2026
View Repository →

real-time-ad-analytics

Kafka + PySpark pipeline for real-time event ingestion, processing, and scalable analytics.

KafkaPySparkDocker
Python • Updated Feb 2026
View Repository →

Online-Voting-System

Backend-driven voting system with authentication, vote handling, and result processing.

FlaskAuthSQL
ASP.NET • Updated Mar 2025
View Repository →

Lie-Detection-System

ML classification pipeline using feature engineering and behavioral data for prediction.

MLFeature Engineering
C++ • Updated Sep 2024
View Repository →

Dog-Breed-Prediction-System

CNN-based image classification system with Flask API for real-time predictions.

TensorFlowCNNFlask
Jupyter Notebook • Updated Sep 2024
View Repository →

Contact

Let’s Connect

I’m actively interested in opportunities across data engineering, backend systems, AI infrastructure, and real-time analytics.