Chirag Mukhija

Lane 04Bengaluru, IndiaBMSCE · CSE ’28

ChiragMukhija

RoleBackend & Systems Engineer

I find what’s actually broken in a system — then fix it, or build it again from the ground up.

Currently building a payment gateway from scratch.B.E. Computer Science & Engineering · CGPA 9.2

025 mHow I work

Find it. Understand it. Fix it — or build it again.

  1. Find the real problem

    Most bugs are symptoms. I trace a failure back to the assumption that broke — two requests racing on one key, a job stranded in “processing”, a query quietly scanning the whole table.

  2. Understand the system

    Before I change anything I map how data moves: who writes, who reads, what has to be atomic, and what is allowed to be eventually right.

  3. Fix it, or build it again

    Then I repair the part that’s wrong — or rebuild the system from zero, in phases, with every trade-off written down where the next engineer can find it.

050 mAbout

Built on discipline.

Portrait of Chirag Mukhija
Sirsa, HRBengaluru, KA

1,860 km as the crow flies

I’m Chirag — a pre-final-year computer science student at BMSCE, Bengaluru, originally from Sirsa, Haryana.

I learn by building the real thing. When a concept matters — idempotency, fan-out, at-least-once delivery — I build a system around it and write down every decision I made, and what it cost.

When I commit to learning something, it’s only a matter of time before I’m as good at it as anyone. Perfection takes longer. I’m fine with that.

Studying
B.E. Computer Science & Engineering
B.M.S. College of Engineering
CGPA
9.2
Years
2024 – 2028
Focus
Backend & distributed systems

Discipline and hard work. Everything else is just the output.

— the one thing I believe in
  1. Rule 01

    Show up on the boring days.

    An hour and a half in the gym every day, around a full college schedule. On LeetCode: 81 active days this year, and a 28-day streak.

  2. Rule 02

    Go deep, not wide.

    40 dynamic-programming problems on LeetCode — Edit Distance, Burst Balloons, Minimum Cost to Cut a Stick — instead of a hundred easy ones.

  3. Rule 03

    Write the decision down.

    Every system I build comes with its trade-offs and known limits written out, so nobody has to guess why it works the way it does.

075 mSystems I built

The deep end.

Three backends built from zero to learn what production systems actually have to get right. Each one is built in phases, and every non-obvious decision is written down — including what it costs.

No “pay now → success” demos. The depth is in the code, so that’s what’s on show.

SYS—01Job-Scheduler

Distributed job scheduler

A task queue built from scratch — the machinery BullMQ, Sidekiq and Celery run under the hood.

7 / 7 phases · complete

  • Node.js
  • PostgreSQL
  • Redis sorted sets
  • Redis Pub/Sub
  • WebSockets
  • Docker Compose
Read the code
Queue simulation · time ×4idle
pending 5retried 0dead 1
waiting for the first event on job:events…

The problem

Background jobs fail, workers die mid-task, and two workers must never run the same job at the same time. The goal: at-least-once delivery with retries, priorities and delayed jobs — without a job ever being silently lost.

Decisions & trade-offs

  1. Postgres is the truth. Redis is only an index.

    Instead of Using Redis as the queue itself

    Every job lives durably in Postgres. Redis holds one sorted set per priority tier, scored by run-at time, so “what’s due next?” is an in-memory lookup — not a table scan.

    Cost: Every path that makes a job eligible has to write to both stores. Forgetting one is the easiest bug to introduce in this codebase.

  2. Claim with a guarded UPDATE, not row locks.

    Instead of SELECT … FOR UPDATE SKIP LOCKED

    Once Redis has picked a single candidate there is nothing left to lock against. UPDATE … WHERE status = 'pending' is the guard; zero rows back means a stale entry, so drop it and try again.

    Cost: The claim path now has a hard dependency on Redis — by design, there is no fallback to scanning the table.

  3. Shut down gracefully, don’t hope.

    Instead of Letting the process die on SIGTERM

    The worker awaits its in-flight job before exiting, so scaling down never strands work. The API closes every WebSocket before server.close() — an open socket blocks it forever, which testing caught.

    Cost: Scaling down now takes as long as the slowest in-flight job — shutdown waits for it.

The crux, in code

api/services/queueService.js
async function claimNextJob() {
  const redis = await getRedisClient();

  for (let tier = MIN_TIER; tier <= MAX_TIER; tier++) {
    while (true) {
      const jobId = await redisQueue.peekDueJob(redis, tier);
      if (!jobId) break;

      const claimed = await pool.query(
        `UPDATE jobs
         SET status = 'processing', attempts = attempts + 1
         WHERE id = $1 AND status = 'pending'
         RETURNING *`,
        [jobId]
      );
      await redisQueue.removeJob(redis, tier, jobId);

      if (claimed.rows.length > 0) return claimed.rows[0];
      // stale Redis entry — already claimed elsewhere, try again
    }
  }
  return null;
}
Redis picks the candidate; Postgres decides who actually gets it.

What breaks — honestly

  • No lease or visibility timeout yet — a SIGKILL’d worker strands its job in “processing”.
  • Exponential backoff (5 s · 2ⁿ) has no jitter and no cap, so mass failures retry in lockstep.
  • No priority aging: a steady flood of priority-0 jobs can starve lower tiers.
  • At-least-once, not exactly-once — job handlers have to be idempotent.

Build log

  1. Postgres-backed queue + submission API
  2. Separate worker process + handler registry
  3. Retries, exponential backoff, dead-letter queue
  4. Priorities + delayed jobs
  5. Redis sorted-set hot claim path
  6. Live dashboard over WebSockets + Pub/Sub
  7. Docker, N workers, graceful shutdown

SYS—02InstaBackend

Instagram-style backend

The backend of a photo-sharing app — auth, uploads, a social graph, and feeds that stay fast deep into the scroll.

7 / 7 phases · complete

  • Node.js
  • PostgreSQL
  • Redis
  • AWS S3 presigned URLs
  • Nginx
  • Docker
Read the code
Feed fan-out · 8 followersfan-out on write
feed_itemsfollowsfollower_id → followee_idpostsORDER BY created_atpostauthoropens feed
Cost per post
8 inserts
one feed_items row per follower
Cost per feed open
1 indexed lookup
WHERE user_id = $1 on a precomputed feed

The problem

A feed is the same data read far more often than it’s written. Do you pay at write time — copy each post into every follower’s feed — or at read time, joining the follow graph live on every open?

Decisions & trade-offs

  1. Build both fan-out strategies, side by side.

    Instead of Picking one from a blog post

    Fan-out-on-write precomputes a feed_items row per follower, so reads are a single indexed lookup. Fan-out-on-read writes once and joins follows × posts on every request. Both are exposed so they can be compared directly.

    Cost: Write fan-out grows with follower count — the celebrity problem — and it currently runs inside the post request.

  2. Keyset pagination everywhere. Never OFFSET.

    Instead of LIMIT … OFFSET …

    OFFSET makes the database walk and throw away every skipped row, and pages drift when new posts arrive. An opaque (created_at, id) cursor stays O(page) at any depth.

    Cost: No “jump to page 40”. Clients can only move forward from a cursor.

  3. The app server never touches image bytes.

    Instead of Streaming uploads through Express

    Clients upload straight to S3 with a presigned URL. The server only signs the request — and before saving a post, checks that the image URL points at its own bucket.

    Cost: Orphaned uploads — signed but never attached to a post — need cleaning up separately.

The crux, in code

src/services/feed.service.js
const fanOutOnWrite = async (post) => {
  const followerIds = await followService.getFollowerIds(post.user_id);
  if (followerIds.length === 0) return;

  const values = [];
  const placeholders = followerIds.map((followerId, i) => {
    values.push(followerId, post.id);
    return `($${i * 2 + 1}, $${i * 2 + 2})`;
  });

  await db.query(
    `INSERT INTO feed_items (user_id, post_id)
     VALUES ${placeholders.join(', ')}
     ON CONFLICT (user_id, post_id) DO NOTHING`,
    values
  );
};
Fan-out-on-write: one post becomes a row in every follower’s feed, idempotently.

What breaks — honestly

  • Fan-out runs synchronously in the post request — a big account makes posting slow. The fix is a queue (see SYS—01).
  • Counts are live COUNT(*) rather than counters: right at this scale, with the point to revisit written down.
  • Notifications are poll-based; there are no WebSockets here.

Build log

  1. JWT access + refresh, Redis-revocable sessions
  2. Profiles, posts, S3 presigned uploads
  3. Follow graph
  4. Feeds: fan-out on write and on read
  5. Likes, comments, notifications
  6. Keyset pagination, search, explore
  7. Docker, Nginx, two-layer rate limiting

SYS—03PaytmentGateway

Payment gateway

A payment API where the hard part isn’t taking money. It’s never taking it twice.

Phase 1 of 6 · in progress

  • Node.js
  • Express 5
  • PostgreSQL
  • raw SQL (pg)
  • next: Redis + BullMQ
Read the code
Idempotency · First requestauto-play
  1. POST /paymentsIdempotency-Key: order_1
  2. SELECT … WHERE (merchant_id, key)miss
  3. BEGIN · INSERT payment · INSERT event · COMMITone connection
  4. 201 Createdpay_7f3a

paymentspayment_events · 1 row

ididempotency_keyamountstatus
pay_7f3aorder_1499.00INITIATED

Payment and its first audit event commit together — or not at all.

The problem

Networks time out and clients retry. A retried payment request has to return the original payment — not quietly create a second charge. And every state change has to leave a trail you can reconcile against a bank later.

Decisions & trade-offs

  1. Idempotency keys are unique per merchant.

    Instead of A globally unique key

    UNIQUE (merchant_id, idempotency_key): two unrelated merchants can both send “order_1”. A SELECT pre-check handles the common retry; the constraint’s 23505 violation is the backstop for the race.

    Cost: Two requests in the same millisecond can both do work before one is rejected — closed properly by a Redis lock in phase 3.

  2. A payment and its first event commit together.

    Instead of Two independent pool.query() calls

    pool.query() can hand BEGIN and INSERT to different connections, which makes the transaction meaningless. One checked-out client runs BEGIN → INSERT payment → INSERT event → COMMIT.

    Cost: Holding a connection for the whole transaction — the pool is capped at 20.

  3. History is append-only. Money is DECIMAL.

    Instead of Overwriting a status column · FLOAT amounts

    payment_events keeps every transition, which is what debugging and reconciliation actually read. DECIMAL(12,2) because binary floats drift on money.

    Cost: A denormalised status column still has to be kept in sync for fast queries.

The crux, in code

src/services/paymentService.js
const client = await pool.connect();
try {
  await client.query('BEGIN');
  const { rows } = await client.query(
    `INSERT INTO payments (merchant_id, idempotency_key, amount)
     VALUES ($1, $2, $3) RETURNING *`,
    [merchantId, idempotencyKey, amount]
  );
  await client.query(
    `INSERT INTO payment_events (payment_id, from_status, to_status)
     VALUES ($1, NULL, $2)`,
    [rows[0].id, rows[0].payment_status]
  );
  await client.query('COMMIT');
  return rows[0];
} catch (err) {
  await client.query('ROLLBACK');
  if (err.code === '23505') throw conflict(409); // lost the idempotency race
  throw err;
} finally {
  client.release();
}
One connection, one transaction — and a clean 409 when the race is lost.

What breaks — honestly

  • The idempotency race isn’t fully closed yet — a Redis lock lands in phase 3.
  • A DB write and a bank call can never be atomic. The honest fix is a reconciliation job (phase 6).
  • No test suite yet.

Build log

  1. Core API, idempotency, transactions, audit trail
  2. Redis locking + retries on bank calls
  3. Per-merchant rate limiting + caching
  4. Docker + Nginx
  5. Reconciliation batch job

100 mProblems I attacked

Other problems I went after.

PRB—0101 / 03

PRactice

LeetCode for open source · team of four

The problem

Beginners can’t practise real open-source contribution: real repositories are too big to read, and far too big to hand to an LLM.

My part

  • Built the Issue → PR → Diff → File mapping that pulls only the code a fix actually touched, instead of stuffing a whole repo into context.
  • Wired the GitHub REST pipeline to fetch pre-fix file contents by commit SHA.
  • Integrated a Monaco editor that reveals the codebase progressively — find the file first, then the code.
sent to the model, not the repo
1 file
sent to the model, not the repo
contributors
4
contributors
  • Next.js
  • TypeScript
  • FastAPI
  • GitHub REST API
  • Monaco
Repository

PRB—0202 / 03

ClipForge

Topic in, finished short out · solo

The problem

One good 60-second explainer takes hours of scripting, voice-over, editing and timing — and every video is built by hand from scratch.

What I did

  • Split the job into 10 durable steps — angle, script, critic, storyboard, voice, assets, timeline, render, metadata, publish — each one a retryable job in SQLite that resumes after a crash.
  • The LLM never writes animation code. It fills props for eight hand-built Remotion templates, so a render can’t break.
  • Animations are anchored to spoken words, not timestamps — change the voice and every cue re-syncs itself.
topic → 1080×1920 MP4
~5 min
topic → 1080×1920 MP4
resumable pipeline steps
10
resumable pipeline steps
  • TypeScript
  • Hono
  • SQLite
  • Remotion
  • ffmpeg
  • Claude · Gemini · OpenAI
Repository

ALSO BUILT03 / 03

Smaller laps.

  • Signal

    AWS hackathon · voice notes → tasks, reminders and calendar events. React Native, Express, Whisper, Gemini.

  • Wanderlust

    Full-stack rental marketplace · auth, role-based authorisation, listings CRUD. MongoDB, Express, Node.

Everything else on GitHub

125 mProblem solving

Depth over count.

I’d rather spend an evening on one hard problem than on ten easy ones. On LeetCode, 76 of my 97 solves are Medium or Hard; 40 are dynamic programming.

Most of my practice happens locally, in Java, one folder per topic — graphs, DP, heaps, tries. That’s where the other hundred-odd problems live.

200+

problems solved across platforms

76

of 97 LeetCode solves are Medium or Hard

1,460

LeetCode contest rating

15 Hard61 Medium21 EasyLeetCode profile

Problems I’ve solved

  • Burst Balloons
  • Edit Distance
  • Minimum Cost to Cut a Stick
  • Partition Array for Maximum Sum
  • Matrix-Chain Multiplication
  • Longest Increasing Subsequence
  • Longest Common Subsequence
  • 0/1 Knapsack
  • Rod Cutting
  • Catalan Numbers
  • Longest Path in a Matrix
  • Target Sum Subset
  • Kosaraju’s SCC
  • Tarjan’s Algorithm
  • Disjoint Set Union
  • Topological Sort
  • Alien Dictionary
  • Prim’s MST
  • Connecting Cities
  • BFS / DFS
  • Tries
  • Segment Trees
  • Sliding Window Maximum
  • K Weakest Rows

150 mAlways in training

Still in training.

  1. 01

    Sigma — Full-Stack Web Development + DSA

    Apna College

    The foundation: JavaScript, Node, MongoDB, React — and DSA in Java.

    Completed
  2. 02

    Fundamentals of Networking for Effective Backends

    Hussein Nasser · Udemy

    TCP vs UDP, NAT, proxies, TLS. Wrote a UDP server and port-forwarding experiments alongside.

    Completed
  3. 03

    Fundamentals of Backend Engineering

    Hussein Nasser · Udemy

    Protocols, communication patterns and execution models — the theory under the systems above.

    In progress

175 mLife

Off the clock.

Swimming

Swimmer first.

I swam competitively in district tournaments and came home with silver in the 100 m freestyle and the 50 m backstroke. Swimming taught me what I still believe: race day only shows the training you’ve already done.

  1. 100 m Freestyle

    District · Silver

    2ND
  2. 50 m Backstroke

    District · Silver

    2ND

Training

Gym, every day.

An hour and a half, seven days a week, around a full college schedule.

Badminton

Badminton on weekends.

Any sport, really — as long as someone’s keeping score.

Music

Music, always.

I love music, and I sing whenever I get the chance.

Reading

Reading.

Self-improvement books, mostly. They all end up saying the same thing: discipline compounds.

200 mThe wall

Touch the wall.

I’m looking for backend and systems internships. If you’re building something that has to hold up under real load, I’d like to hear about it.

chiragmukhija.cs24@bmsce.ac.in