Head of Infrastructure
San FranciscoRemoteDirectorFULL_TIMEtoday
General Compute is the neocloud for alternative chips. Inference is fragmenting: purpose built silicon from SambaNova, Cerebras, Positron, d Matrix, and others already beats GPUs on decode, and we productionize that hardware — we buy the racks, find the data center space, and run it for our customers.
About Us
General Compute is the neocloud for alternative chips.
Inference is fragmenting: purpose-built silicon from SambaNova, Cerebras, Positron, d-Matrix, and others already beats GPUs on decode, and we productionize that hardware — we buy the racks, find the data center space, and run it for our customers. Each piece of hardware runs the workload it's actually built for: prefill stays on GPUs, decode moves to the chip built for it, and today that means generating tokens 5–7× faster than existing GPU-based competitors. Our customers are frontier labs, fast-growing AI application companies, and asset-light clouds.
We closed a $15M seed round in May 2026, and have since closed a $400M debt facility — $100M funded upfront by Upper90, with the balance available for drawdown — collateralized by our inference chips.
About the role
You'll own the infrastructure layer of our inference cloud end-to-end. Today that means the control plane, the gateway in front of our ASIC fleet, and the observability stack that tells us where every millisecond goes. Over the next 6-8 months, it will grow into a heterogeneous fleet: ASICs for decode, GPUs for pre-fill, and the physical-layer ownership that comes with it.
The first six months are hands-on: k8s manifests, dashboards, oncall, and a direct line to our ASIC partner's engineering team when production behaves strangely. The team grows under you from there.
What you'll do:
- Own the inference control plane. At the moment, it's built on configuration provided by our ASIC partner; you'll be the person who understands it deeply enough to modify, extend, and eventually replace pieces of it.
- Own the gateway and load balancer that fronts the fleet. Model placement, request routing, and tail-latency engineering live here, driven by live utilization and per-model SLOs.
- Own observability end-to-end. Per-request tracing from OpenRouter ingress through to the accelerator, with p50/p95/p99 dashboards, SLOs, and alerting that wakes the right person.
- Run capacity planning against a real, distributed traffic mix across the open-weight models we serve.
- Own the operational side of the ASIC partnership. Most weird production issues route through their engineering team until we build that expertise in-house, and you'll be our technical face in those conversations.
- Bring up the pre-fill side of our disaggregated architecture on a second hardware platform as it comes online. Different vendor, different fabric, different kernels.
- Build the on-call and incident response practice from zero. Hire and grow the team underneath you.
What we need from you:
- 7+ years in infrastructure, SRE, or platform engineering, with at least some of it at a serious inference, ML, or HPC shop.
- Hands-on with Kubernetes at production scale — not just deploying, but debugging the weird stuff.
- Strong instincts for tail latency. You think about p99 and utilization as the same problem, not different ones.
- Comfortable owning a vendor relationship where the vendor's bugs are now your production issues.
- Track record of building observability practices that actually catch problems, not just generate dashboards.
- Have been on-call through real incidents and can talk about what you learned.
- Want to be the first infra hire at something early, not the tenth at something big.
NIce-to-Haves:
- Experience operating non-NVIDIA accelerators in production — TPUs, ASICs, or alternative GPU vendors.
- Background with model-serving stacks (vLLM, TGI, TensorRT-LLM, SGLang).
- Network fabric experience at data-center scale (RoCE, InfiniBand).
- Have hired and managed an infra team before.
- Comfort at the hardware boundary — firmware, drivers, thermals — for when the roadmap takes us there.
Frequently asked
Is this Head of Infrastructure role remote?
Yes — this role is remote-friendly (San Francisco).