A pipeline that converts scanned engineering diagrams into structured relational graphs for industrial data analytics, deployed behind a containerized API for symbol classification at scale.
Software engineer whose real strength is turning messy, unstructured input — scanned diagrams, scraped web pages, live sensor feeds — into structured, queryable systems. Built for an audience evaluating data-pipeline fundamentals, not a claim to warehouse-scale production tooling I haven't used yet.
Dual-degree engineering student at IIT Kharagpur (B.Tech Hons. + M.Tech, graduating July 2026) with roughly two years of hands-on build time across three internships and a run of independent projects. The common thread across that work is data transformation: taking raw, inconsistent, real-world input and shipping a pipeline that turns it into something structured, monitored, and usable downstream — the same problem ETL work solves, even where the specific tools differ.
Each of these moved raw, messy input through an automated pipeline into a structured, deployed output.
A pipeline that converts scanned engineering diagrams into structured relational graphs for industrial data analytics, deployed behind a containerized API for symbol classification at scale.
An ingestion pipeline that scrapes source content, embeds it, and serves it back through a production microservice for semantic retrieval.
Turned raw equipment and geological sensor readings into monitored dashboards and predictive alerts for an operations team.
Event-driven matching: incoming blood requests trigger real-time proximity lookups against a live donor dataset and dispatch alerts across channels.
Research prototype: mine-site sensor readings are checked against compliance thresholds on-chain, and only compliant readings mint a certificate — with a backend that turns those on-chain events back into structured, mapped data.
Stylized low-poly reconstruction of the project's sensor breadboard rig (Arduino, gas sensors, LCD readout) — simplified for the web, not to exact scale. Drag to rotate.
Grouped by what's actually production-tested versus what's foundational or in progress.
This role's core stack — Airflow / Prefect, dbt, and cloud warehouses like Redshift or Snowflake — isn't something I've shipped with yet. My pipeline work above covers the same underlying problem (structuring messy input, monitoring pipeline health, deploying reliably) using a different toolset. Flagging that directly rather than overstating it.