Gridcast: NYISO day-ahead load forecasting
Personal project · 2026
Tomorrow's hourly electricity load for New York's 11 grid zones and the statewide total, forecast
before the day-ahead bid cutoff and scored every day against what actually happened. NYISO publishes its own
day-ahead forecast, which makes a hard, public baseline.
What I built
- A global Temporal Fusion Transformer trained across all 11 zones and the statewide total on cutoff-aligned
data, so the model only sees what would exist before the bid cutoff.
- Served as ONNX through FastAPI on GCP, with daily ingestion, self-scoring against NYISO and a LightGBM
fallback.
- Deployed to Kubernetes: a kind cluster in CI on every push, and a 72-hour GKE Autopilot trial with a
digest-pinned model image.
How it measures up
- On par with NYISO's own day-ahead forecast over a 12-month backtest: 4.80% vs. 4.85%
average hourly error (MAPE), no significant difference under a paired bootstrap.
- 53% lower error than the best naive baseline.
- 20+ controlled experiments on features, history and models. Conformal calibration raised how often actual
load falls inside LightGBM's 80% forecast range from 57% to 77%, still short of the
80% target.
Decisions
- Fixed the reporting rules before the first result.
- Rejected a weather feed that scored better, because part of it arrives after the bid cutoff: in live use
it would not exist in time.
Personal project · 2026
A voice mock-interview app for ML and AI engineering interviews, grounded in rubrics and tailored
to each user's résumé and target job. Grading every answer with a frontier model is expensive, so the core
question was which grader to trust.
What I built
- The voice interview flow, rubric-grounded feedback, and hybrid BM25 + BGE retrieval: how often the right
passage ranks in the top 5 (Recall@5) rose from 72% to 84% on 61 paraphrased queries.
- A deployed app with rate limits and usage budgets.
The grader study
- 598 Claude-labeled synthetic answers, question-grouped splits with about 20% held out per seed, five
seeds.
- A LoRA-tuned Qwen3-4B agreed with Claude's grades at QWK 0.94
(1.0 = perfect agreement), against 0.79 for the best baseline, and beat the scikit-learn graders on every
seed.
- vLLM serves a grading call in 30 ms (p95).
Decisions
- Hosted grading runs on DeepSeek: 94% of its grades fall within one point of Claude's, at about
1/200th of the cost.
- But DeepSeek gave the same grade twice only 57% of the time, so it never makes training labels.
AI Tender Evaluation Assistant
Team project with a second engineer · Techlent Machine Learning Engineer Fellowship (training program) · 2026 · real-bid evaluation in October 2026
Bids run to hundreds of pages, often scanned. The design: a vision model reads each form, each
value is verified on the PDF's text layer or by an independent second read, and a deterministic rule engine
decides pass, review or disqualified, with each value citing its page. I owned the FastAPI backend, the LLM
gateway, the bid checks and the evaluation; my teammate built the document parser, the form templates and the
original review UI.
What I built
- Verification for every extracted value: values the two reads disagree on go to a reviewer.
- An LLM gateway that enforces data handling in code: confidential projects reach only local models, and
redacted samples reach only one cleared Azure endpoint, with a cache, a per-project daily budget and a rate
limit.
- Postgres-backed jobs with checkpoints, approval pauses and rate-limit backoff, and a bounded search agent
tested against prompt injection. Re-checking 10 bidders after a rule-set edit takes
0 model calls.
- The evaluation harness: one answer key per bid (a model prefilled it from the page images, I checked every
value and froze it with a SHA-256), a scorer that prints numbers only, and diagnostics that show a request's
shape (field names, counts, page numbers) but never document content.
Results on two real bids
- Two real tenders, each a 62-page scanned offer, about 70 scored values in all. Tender 1 was the development
bid; Tender 2 was held out and run once, at a recorded commit, with the local model.
- 14 of 14 prices, quantities and dates read correctly, on both bids, with all three
models tried.
- Held-out bid: 13 of 13 items found present or absent,
no wrongful disqualification, and 21 of 35 values exact overall: numbers and dates
5 of 5, with long free text the weak part. The bid check took 38 min 47 s, at $0.
- Every value verification flagged on the held-out bid was indeed wrong (8 of 8), but it also accepted 8 wrong
values out of 29, which is why a person reviews.
- Values cited on their own page, on multi-page forms: 5 of 23 → 21 of 23 after a
separate locate step, developed on the development bid and measured on the held-out bid (after the fixes).
- Same pipeline and keys on three models. Values exact: 23/38 and 21/35 for the local Qwen3-VL-8B, 20/38 and
18/35 for gpt-4.1-mini, 21/37 and 18/34 for gpt-5-mini. Per tender: $0 and about 55 min locally, about $0.27
and 6-9 min on gpt-4.1-mini, about $0.43 and 25 min on gpt-5-mini. The reasoning model drafted fuller rule
sets (about 25% fewer gaps) but sent many more checks to a person.
Decisions
- Rehearsed the whole local pipeline on a synthetic scanned bid first (15/15 items, 16/16 values, nothing
invented). It caught a reasoning-model runaway and an 8 GB hidden prompt cache before any real document was
used.
- Fit the run on a 16 GB Mac mini: the instruct build of Qwen3-VL instead of a thinking build (those wrote
about 2,400 tokens of reasoning before a one-line answer), a dedicated Ollama server with the prompt cache
off, and a 32k context with an 8-bit KV cache at the memory cost of 16k.
- Froze the commit and settings before the held-out bid and ran it once; later re-runs are labelled "after
the fixes".
- Fixed all 11 defects the scorer exposed, each with a test. Two close paths to a false disqualification:
switching to a cloud model disqualified an offer over two blanks, because it answers "" where the local model
says null and the blank-safety rule only checked for null; and a form found on its pages but read as absent
now goes to review instead. One was my own regression, a fix that made the model mark 19 of 28 read fields
as blacked out; I measured it and fixed it.
- Chose the job runner with a 12-scenario fault-injection harness: Procrastinate and LangGraph both passed
all 7 must-pass gates, so the simpler Postgres queue won, with about 5 s recovery
and 0 lost jobs.
- Shipped on Azure Container Apps with Entra ID sign-in, Key Vault and OIDC deploys that store no secrets,
plus a public no-backend demo; audited the 455-commit history for the public release (AGPL-3.0). About 750
Python, 19 Postgres, 119 Vitest and 5 Playwright tests run in CI.
The bids are redacted sample documents provided by the training program and cleared for
cloud-model testing; there is no client. Two bids is a small sample: the numbers show the method, not a
production rate, and no bid was decided without a person.