Measure whether a training or serving change improved real throughput.
Mohamed A M Elansary, PhD — scientific/HPC execution, production systems, and honest measurement for RL training and inference loops.
Scientific compute
- Six-plus years of multimodel, multi-basin forecast experiments across hydroclimates on Linux/HPC (UCAR, TACC Stampede).
- Compared statistical and physically based stacks, quantified uncertainty, and reported regime-dependent failure modes rather than a single flattering score.
- That is the measurement analogue of asking whether a data-generation or training-loop change improved end-to-end throughput.
Production systems
- Production GPT, Claude, and Gemini agent workflows with retrieval, routing, tenant isolation, provenance, and regression evaluation sets at Vertexium, including a multi-tenant conversational receptionist.
- That maps to production systems sitting next to a serving stack. It is not Harmonic RL training infrastructure, CUDA kernels, or Aristotle/Lean work.
- A PhD in a related field is a posted preferred qualification; this profile includes an Environmental Engineering PhD.
Proposed first contribution
For one RL training or inference loop already in flight, define what peak performance and optimal throughput mean versus a local score that is easy to move. Write a small failure taxonomy: a kernel or scheduler change that does not move end-to-end RL step time; GPU idle while CPU Lean/REPL work waits; a sharding choice that hits memory limits on a multi-billion-parameter cycle; a serving change that helps a microbenchmark and hurts production latency. Stand up a small measurement with provenance on traces, compare simple baselines, attach uncertainty, and write a clear report before expanding kernel, router, or sharding work. This is a proposed measurement approach, not a claim of prior Harmonic-internal work, CUDA authorship, invented metrics, safety research, or RLHF.
Honest fit boundary
Large-scale training/inference engineering and Harmonic's math/formal-reasoning product stack is a stretch. I have not owned proprietary RL training or serving infrastructure, written CUDA or Triton kernels, scaled models with FSDP or tensor parallelism on multi-node GPU clusters, deployed performant LLM inference engines, worked in Lean 4 or on Aristotle, or claimed Harmonic-internal work, and I do not invent metrics, safety research, or RLHF. The credible contribution is scientific/HPC execution, production systems, data pipelines, and honest measurement.
Role and location
Research Engineer, Training & Inference · Palo Alto · OnSite. Ashby lists workplaceType OnSite and isRemote false, with address locality Palo Alto, CA, United States. Willing to relocate to Palo Alto with a relocation package. Remote work is not asserted.
Posting compensation: “$200K – $450K • Offers Equity”. · Official role posting