← cyfang.org

Chenyu Fang, Ph.D.

Machine learning & AI engineer · Physics Ph.D.

I take models from data to deployment and hold them to an honest baseline: time-series forecasting, rare-event classification, and LLM applications with the evaluation that keeps them honest. Seven years of Ph.D. research finding rare signals in simulated particle-physics data taught me to fix the rules before looking at results; since 2026 I have applied that to forecasting, LLM grading and document AI.

Projects

Gridcast: NYISO day-ahead load forecasting

Personal project · 2026

Tomorrow's hourly electricity load for New York's 11 grid zones and the statewide total, forecast before the day-ahead bid cutoff and scored every day against what actually happened. NYISO publishes its own day-ahead forecast, which makes a hard, public baseline.

What I built

  • A global Temporal Fusion Transformer trained across all 11 zones and the statewide total on cutoff-aligned data, so the model only sees what would exist before the bid cutoff.
  • Served as ONNX through FastAPI on GCP, with daily ingestion, self-scoring against NYISO and a LightGBM fallback.
  • Deployed to Kubernetes: a kind cluster in CI on every push, and a 72-hour GKE Autopilot trial with a digest-pinned model image.

How it measures up

  • On par with NYISO's own day-ahead forecast over a 12-month backtest: 4.80% vs. 4.85% average hourly error (MAPE), no significant difference under a paired bootstrap.
  • 53% lower error than the best naive baseline.
  • 20+ controlled experiments on features, history and models. Conformal calibration raised how often actual load falls inside LightGBM's 80% forecast range from 57% to 77%, still short of the 80% target.

Decisions

  • Fixed the reporting rules before the first result.
  • Rejected a weather feed that scored better, because part of it arrives after the bid cutoff: in live use it would not exist in time.
  • PyTorch
  • Temporal Fusion Transformer
  • LightGBM
  • conformal prediction
  • ONNX
  • FastAPI
  • GCP
  • Kubernetes

MLE/AIE Interview Coach

Personal project · 2026

A voice mock-interview app for ML and AI engineering interviews, grounded in rubrics and tailored to each user's résumé and target job. Grading every answer with a frontier model is expensive, so the core question was which grader to trust.

What I built

  • The voice interview flow, rubric-grounded feedback, and hybrid BM25 + BGE retrieval: how often the right passage ranks in the top 5 (Recall@5) rose from 72% to 84% on 61 paraphrased queries.
  • A deployed app with rate limits and usage budgets.

The grader study

  • 598 Claude-labeled synthetic answers, question-grouped splits with about 20% held out per seed, five seeds.
  • A LoRA-tuned Qwen3-4B agreed with Claude's grades at QWK 0.94 (1.0 = perfect agreement), against 0.79 for the best baseline, and beat the scikit-learn graders on every seed.
  • vLLM serves a grading call in 30 ms (p95).

Decisions

  • Hosted grading runs on DeepSeek: 94% of its grades fall within one point of Claude's, at about 1/200th of the cost.
  • But DeepSeek gave the same grade twice only 57% of the time, so it never makes training labels.
  • LLM evaluation
  • LoRA / PEFT
  • Qwen3-4B
  • vLLM
  • RAG
  • BM25 + BGE
  • FastAPI

AI Tender Evaluation Assistant

Team project with a second engineer · Techlent Machine Learning Engineer Fellowship (training program) · 2026 · real-bid evaluation in October 2026

Bids run to hundreds of pages, often scanned. The design: a vision model reads each form, each value is verified on the PDF's text layer or by an independent second read, and a deterministic rule engine decides pass, review or disqualified, with each value citing its page. I owned the FastAPI backend, the LLM gateway, the bid checks and the evaluation; my teammate built the document parser, the form templates and the original review UI.

What I built

  • Verification for every extracted value: values the two reads disagree on go to a reviewer.
  • An LLM gateway that enforces data handling in code: confidential projects reach only local models, and redacted samples reach only one cleared Azure endpoint, with a cache, a per-project daily budget and a rate limit.
  • Postgres-backed jobs with checkpoints, approval pauses and rate-limit backoff, and a bounded search agent tested against prompt injection. Re-checking 10 bidders after a rule-set edit takes 0 model calls.
  • The evaluation harness: one answer key per bid (a model prefilled it from the page images, I checked every value and froze it with a SHA-256), a scorer that prints numbers only, and diagnostics that show a request's shape (field names, counts, page numbers) but never document content.

Results on two real bids

  • Two real tenders, each a 62-page scanned offer, about 70 scored values in all. Tender 1 was the development bid; Tender 2 was held out and run once, at a recorded commit, with the local model.
  • 14 of 14 prices, quantities and dates read correctly, on both bids, with all three models tried.
  • Held-out bid: 13 of 13 items found present or absent, no wrongful disqualification, and 21 of 35 values exact overall: numbers and dates 5 of 5, with long free text the weak part. The bid check took 38 min 47 s, at $0.
  • Every value verification flagged on the held-out bid was indeed wrong (8 of 8), but it also accepted 8 wrong values out of 29, which is why a person reviews.
  • Values cited on their own page, on multi-page forms: 5 of 23 → 21 of 23 after a separate locate step, developed on the development bid and measured on the held-out bid (after the fixes).
  • Same pipeline and keys on three models. Values exact: 23/38 and 21/35 for the local Qwen3-VL-8B, 20/38 and 18/35 for gpt-4.1-mini, 21/37 and 18/34 for gpt-5-mini. Per tender: $0 and about 55 min locally, about $0.27 and 6-9 min on gpt-4.1-mini, about $0.43 and 25 min on gpt-5-mini. The reasoning model drafted fuller rule sets (about 25% fewer gaps) but sent many more checks to a person.

Decisions

  • Rehearsed the whole local pipeline on a synthetic scanned bid first (15/15 items, 16/16 values, nothing invented). It caught a reasoning-model runaway and an 8 GB hidden prompt cache before any real document was used.
  • Fit the run on a 16 GB Mac mini: the instruct build of Qwen3-VL instead of a thinking build (those wrote about 2,400 tokens of reasoning before a one-line answer), a dedicated Ollama server with the prompt cache off, and a 32k context with an 8-bit KV cache at the memory cost of 16k.
  • Froze the commit and settings before the held-out bid and ran it once; later re-runs are labelled "after the fixes".
  • Fixed all 11 defects the scorer exposed, each with a test. Two close paths to a false disqualification: switching to a cloud model disqualified an offer over two blanks, because it answers "" where the local model says null and the blank-safety rule only checked for null; and a form found on its pages but read as absent now goes to review instead. One was my own regression, a fix that made the model mark 19 of 28 read fields as blacked out; I measured it and fixed it.
  • Chose the job runner with a 12-scenario fault-injection harness: Procrastinate and LangGraph both passed all 7 must-pass gates, so the simpler Postgres queue won, with about 5 s recovery and 0 lost jobs.
  • Shipped on Azure Container Apps with Entra ID sign-in, Key Vault and OIDC deploys that store no secrets, plus a public no-backend demo; audited the 455-commit history for the public release (AGPL-3.0). About 750 Python, 19 Postgres, 119 Vitest and 5 Playwright tests run in CI.

The bids are redacted sample documents provided by the training program and cleared for cloud-model testing; there is no client. Two bids is a small sample: the numbers show the method, not a production rate, and no bid was decided without a person.

  • document AI
  • vision LLMs
  • Qwen3-VL
  • Ollama
  • LLM evaluation
  • guardrails
  • FastAPI
  • PostgreSQL
  • Azure

Research

Rare-signal benchmark: XGBoost vs. LorentzNet vs. Particle Transformer

Research Assistant (part-time), University of Oklahoma · 02/2026 to present

Does a modern network beat gradient boosting at finding a rare signal, and is the gain worth the cost? Extended my published analysis with new background simulations to find out.

  • Replaced random accept/reject tagging with probability weights validated against direct tagging, keeping expected yields unbiased with a 9‑19x larger effective sample.
  • Compared XGBoost, a re-implemented LorentzNet, a Particle Transformer (ParT) and a Deep Sets control on identical splits under rules fixed in advance, three seeds and paired bootstraps.
  • ML beat the cut-based selection by 9‑21% in expected significance, and both networks beat XGBoost's AUC.
  • LorentzNet matched ParT at a tenth of its size with 2x faster inference, while Deep Sets fell below XGBoost: relational structure drives the gain. XGBoost trailed by under 0.01 AUC at about 1/100th the training cost.
  • At the higher mass, limited simulation statistics, not the model, cap the result. Every study ran on a resumable, hash-verified data pipeline.
  • XGBoost
  • graph neural networks
  • transformers
  • importance weighting
  • bootstrap inference

Higgs boson rare-signal detection

Ph.D. research, University of Oklahoma · 08/2018 to 12/2025

  • Designed a simulation-based search for a new particle: produced the collision datasets and tuned the rule-based selection that was published and became the baseline for later ML models.
  • Python/C++ pipelines for multi-million-event simulated datasets; domain-informed features and XGBoost classifiers that separate a rare signal from backgrounds about 107 times larger.
  • Evaluated with sample-weighted ROC/AUC and discovery significance, choosing thresholds on stable validation plateaus rather than noisy peaks. The final model recovers 41.2% of signal at a 7.7% false-positive rate, a projected 6.2σ result.
  • Interpreted with SHAP and feature ablations; packaged as a reproducible CLI with 69 tests, versioned model releases tied to git commits and a 2×10−6 score regression check.
  • Python
  • C++
  • XGBoost
  • SHAP
  • rare-event detection

Experience

Skills

Machine learning
PyTorch, scikit-learn, XGBoost, LightGBM, graph neural networks, transformers (TFT, ParT), time-series forecasting, conformal prediction, SHAP
Methods
Experimental design, bootstrap inference, importance weighting, uncertainty quantification, LLM evaluation
LLM systems
RAG and hybrid retrieval, LoRA / PEFT fine-tuning, distillation, vLLM, local models (Ollama), LangGraph, MCP, guardrails, vision-language models, document AI
Engineering
FastAPI, PostgreSQL / pgvector, job queues, Docker, Kubernetes (GKE, kind), ONNX, GitHub Actions, pytest, Linux, Git
Cloud
GCP (Vertex AI, Cloud Run), Azure (Container Apps, Entra ID, Key Vault), AWS
Languages
Python, C++, SQL

Publications and talks

Parallel talk, PHENO 2025 (particle-physics conference) · Avenir Graduate Fellowship (2023, 2024)

Education