Dinesh Jinjala

07A multi-tenant Pharma Manufacturing Analytics SaaS

Batch Evaluation & Export Engine

Built v1, main developer of v2

I built the engine that runs scientists' calculations across manufacturing batch data, and made it memory-bounded and crash-safe: large jobs stream instead of loading everything, and a crashed worker doesn't lose or duplicate a result. In development benchmarks, streaming export cut memory for 1M+ rows from ~4 GB to ~80 MB.

~4 GB → ~80 MB
memory to export 1M+ rows (development benchmark)
1.8×
faster Excel export (development benchmark)
~3×
faster key query (development benchmark)

Python / PostgreSQL / Redis

Problem

Scientists define calculations that run on every manufacturing batch, using data from several databases. As data grew, large evaluations and exports hit memory limits, and in a GxP system every job must finish exactly once, even when a worker crashes.

Approach

  • I built the first evaluation engine, then did most of the work on its rebuild as a service.
  • Calculations run in dependency order, and data streams through in chunks, so memory stays flat as data grows.
  • Each job has exactly one owner at a time, and recovery after a crash is safe to replay.
  • Exports stream straight to the file instead of loading full result sets.

Architecture

  1. Job table
  2. Claim
  3. Streaming fetch
  4. Formula execution
  5. Safe writes
  6. Results & export

Outcome

Measured in development, exporting 1M+ rows dropped from about 4 GB to about 80 MB of memory, and Excel export became 1.8× faster with half the memory. Large exports no longer run out of memory, and a crashed job recovers to the exact final state.

What I learnedThe biggest memory win came from never holding a full result set.

Building something like this?

Tell me about it