Hi, I'm Sumit.
Backend & Data Infrastructure Engineer
11+ years building large-scale product backends and the data platforms behind them. Twice a founding engineer — I design and own systems end-to-end, from API contracts to ClickHouse clusters, and I'm usually the person holding the pager.
Over 11+ years I've built and run backends handling hundreds to thousands of daily active users, millions of requests a day, terabytes of data, and more — usually as the sole owner.
The short version.
Senior software engineer with 11+ years across product backends and large-scale data infrastructure. I've been the founding engineer twice — at TrulyMadly and JobHai — owning the architecture, the databases, and the pipelines from the first commit.
My core is ClickHouse and the systems around it: Kafka and CDC ingestion, MergeTree storage design, and hybrid hot/cold tiering. Lately I've extended into applied AI, building a production text-to-SQL platform (LiteLLM + Gemini) with a semantic layer over business-critical schemas.
A high-ownership IC — I tend to be the person who designs it, ships it, and keeps it running.
- 11+ years experience
- Noida, India
- Backend · Data · Applied AI
- Open to opportunities
Focus areas.
Backend Architecture
Node.js / Python REST APIs, event-driven systems, and MySQL/Redis at scale — from API contracts to production reliability.
Data Infrastructure
ClickHouse cluster design, Kafka + CDC ingestion pipelines, MergeTree/TTL storage strategy, and hybrid hot/cold tiering.
Recommendation & ML
Collaborative-filtering + behavioural-signal engines and clickstream/feature pipelines powering millions of daily notifications.
Applied AI
Text-to-SQL over production data (LiteLLM + Gemini), semantic layers, and LLM tooling embedded into BI (forked Apache Superset).
The stack.
Where I've worked, and what I owned.
- Sole backend architect for JobHai — designed and built the entire backend, now serving high-throughput production traffic with no dedicated ops team.
- Architect and operate ClickHouse clusters at massive scale — sole owner of schema design, index strategy, and cluster health.
- Built Kafka → ClickHouse pipelines processing 100M+ rows/day, incl. CDC sync from MySQL across 26 tables at ~1-minute latency.
- Multi-datacenter ClickHouse migration with hybrid hot/cold tiering — eliminated a 95% disk-capacity crisis, ~₹2 Cr annual savings.
- Designed and shipped a text-to-SQL AI analytics platform (LiteLLM + Gemini, forked Apache Superset) covering 37 charts and all core metrics.
- Reduced recommendation pipeline runtime from 12h → 1h, enabling 10M+ targeted push notifications daily.
- Greenfield build of JobHai's backend — service architecture, database schema, and API contracts from day one; all still in production.
- Built the candidate/job recommendation engine (collaborative filtering + behavioural signals), contributing to 30% growth in daily active users.
- Built large-scale clickstream ingestion + behavioural analytics that became the foundation for all ML feature generation.
- Owned hiring and technical roadmap for a small sub-team — interviewed, onboarded, and mentored engineers.
- Founding backend engineer — designed and built the entire backend (core APIs, matching logic, data pipelines) as the platform grew to 1M–5M users.
- Built the Recommendation Engine powering match suggestions — hybrid collaborative filtering that improved engagement by 60% and killed cold-start starvation.
- Designed and implemented the full payments backend (subscriptions, premium gating, transaction reporting) — moving the product to break-even.
Selected work.
Most of the systems below are proprietary, so they're described rather than linked. The open-source work underneath is public.
Text-to-SQL Analytics Platform
2024–25 · JobHaiNatural-language query interface over ClickHouse — LiteLLM + Gemini 2.0 Flash with a custom React chat widget embedded in a forked Apache Superset, over a semantic layer encoding every core business metric.
JobHai Backend
2019–present · JobHaiGreenfield backend for a B2B jobs marketplace — REST APIs, MySQL at scale, Redis caching, Kafka pipelines, the recommendation engine, and the analytics stack. Sole architect across every layer.
ClickHouse Multi-DC Migration
2024 · JobHaiMigrated a massive node at 95% capacity to a two-node hybrid cluster with an S3 cold tier — storage policies, native BACKUP/RESTORE strategy, and an upgrade path to ClickHouse 26.x LTS.
MySQL → ClickHouse CDC Pipeline
2023 · JobHaiReal-time change-data-capture across 26 production MySQL tables into ClickHouse using Kafka; the foundation powering all analytics and the ML feature stores.
TrulyMadly Backend
2014–2019 · TrulyMadlyFounding backend for a dating platform — core APIs, the match recommendation engine, the full payments backend (subscriptions, premium gating, reporting), and internal analytics.
MULTI_VALUE type and two-tier operators. Merged into Apache Superset.
Python/TS
MiniCal
A native macOS menu-bar calendar — one click for a clean monthly view. Ships via
its own Homebrew tap.
Swift
cldap
LDAP / Active Directory query CLI — configure once, query anytime.
go install-able.
Go
object-presigner npm
Zero-dependency AWS SigV4 presigned URLs for any S3-compatible storage; separates
the signing host from the CDN host.
JavaScript
// education
// honors & awards
Let's talk.
Building something at scale?
A backend, a data platform, or an AI layer over your data — I'd love to hear about it. Send a message and I'll get back to you.
- thedeceptio@gmail.com
- Noida, India
- github.com/thedeceptio