Portfolio · Data & Backend Engineering

Subham Sarkar

Software engineer whose real strength is turning messy, unstructured input — scanned diagrams, scraped web pages, live sensor feeds — into structured, queryable systems. Built for an audience evaluating data-pipeline fundamentals, not a claim to warehouse-scale production tooling I haven't used yet.

About

Dual-degree engineering student at IIT Kharagpur (B.Tech Hons. + M.Tech, graduating July 2026) with roughly two years of hands-on build time across three internships and a run of independent projects. The common thread across that work is data transformation: taking raw, inconsistent, real-world input and shipping a pipeline that turns it into something structured, monitored, and usable downstream — the same problem ETL work solves, even where the specific tools differ.

Featured Pipelines

Each of these moved raw, messy input through an automated pipeline into a structured, deployed output.

P&ID Diagram Analyzer COE-SEA, IIT Kharagpur · Jan–Apr 2025

A pipeline that converts scanned engineering diagrams into structured relational graphs for industrial data analytics, deployed behind a containerized API for symbol classification at scale.

Input500+ raw P&ID diagram scans (unstructured images)
ProcessPyTorch Geometric GNN for symbol classification & graph construction
OutputStructured relational graphs, queryable for analytics
PythonPyTorch GeometricFlask DockerGitHub Actions CI/CD
Automated CI/CD took this from notebook experiment to a production-deployable classification API.
Chatbot Retrieval Pipeline xLayer Technologies · Aug 2025–Present

An ingestion pipeline that scrapes source content, embeds it, and serves it back through a production microservice for semantic retrieval.

Input200+ scraped pages of unstructured web content
ProcessHugging Face embedding models → Milvus vector store (250+ embeddings)
OutputSemantic search / context retrieval for a live chatbot API
PythonFastAPIMilvus Hugging FaceDockerAWS S3JWT
Runs as a standing microservice, not a one-off batch job — ongoing ingestion, not a single import.
KPI Reporting & Anomaly Detection Elitech Earth Science · Jun–Jul 2024

Turned raw equipment and geological sensor readings into monitored dashboards and predictive alerts for an operations team.

InputRaw equipment & geological sensor readings
ProcessEDA, anomaly detection, predictive modeling
OutputPower BI KPI dashboards + Flask-served predictions
PythonFlaskPower BI
Real-time KPI monitoring cut equipment downtime by 20%; predictive models cut error rates by 12% versus prior statistical methods.
Real-Time Donor Alert System TechAThon Hackathon · Aug 2025

Event-driven matching: incoming blood requests trigger real-time proximity lookups against a live donor dataset and dispatch alerts across channels.

Input1,000+ live donor geolocation entries + incoming requests
ProcessProximity matching, event-triggered dispatch
OutputMulti-channel alerts (Twilio, Firebase, email)
Node.jsExpressNext.js LeafletTwilioFirebase
Built end-to-end, including the real-time, event-driven matching layer, in a 48-hour build window.
Blockchain Compliance Ledger IIT Kharagpur Research · Ongoing

Research prototype: mine-site sensor readings are checked against compliance thresholds on-chain, and only compliant readings mint a certificate — with a backend that turns those on-chain events back into structured, mapped data.

InputCO2 / PM2.5 / SO2 / noise sensor readings
ProcessSolidity threshold check → certificate mint (Hardhat, testnets)
Outputethers.js listener → enriched GeoJSON via Express API
SolidityHardhatethers.js Node.jsExpressPython (RSA signing)
Research/prototype-stage, deployed to Polygon Amoy & Ethereum Sepolia testnets — not a production system.

Stylized low-poly reconstruction of the project's sensor breadboard rig (Arduino, gas sensors, LCD readout) — simplified for the web, not to exact scale. Drag to rotate.

Skills

Grouped by what's actually production-tested versus what's foundational or in progress.

Languages & Data

PythonTypeScript / JavaScript SQL (DBMS fundamentals)C++

Pipelines & ML

PyTorchTensorFlowLangChain MilvusPrismaMongoDB

Infra & Delivery

DockerAWSTerraform KafkaGitHub Actions CI/CD

Actively building toward

This role's core stack — Airflow / Prefect, dbt, and cloud warehouses like Redshift or Snowflake — isn't something I've shipped with yet. My pipeline work above covers the same underlying problem (structuring messy input, monitoring pipeline health, deploying reliably) using a different toolset. Flagging that directly rather than overstating it.

AirflowPrefectdbt Redshift / Snowflake

Education

Indian Institute of Technology Kharagpur
Dual Degree — B.Tech (Hons.) Mining Engineering + M.Tech Safety Engineering · CGPA 7.00/10
2021 – Jul 2026