Senior AI Infrastructure Engineer
Software Engineering, Other Engineering, Data Science · Full-time
Melbourne, VIC, Australia
We’re Heidi.
We're building the future of healthcare by giving every clinician the earth's finest AI Care Partner. In just 18 months, our clinical AI products have absorbed the administrative chaos of 73 million patient visits. Today, we support over 2.5 million patient sessions a week across 190+ countries.
Healthcare systems are failing us; clinicians spend more time on documentation than on patients, and the human connection that makes medicine worth practicing is eroding. Our mission is simple: double the world’s healthcare capacity and strengthen the human connection at its heart.
We found product-market fit with a freemium medical scribe that clinicians love. Now, we're expanding. Every task a clinician hands to Heidi is a patient who feels more attended to, a health system unclogged, and a clinician who gets to be a clinician again.
If you don’t choose easy and you want to build something way bigger than yourself then, choose the challenge, choose Heidi.
The role
This role sits in the model team, the researchers and engineers who train, deploy, and own the AI models behind every Heidi product. You’ll build and operate the infrastructure that makes those models fast, reliable, and cost-effective at scale.
Your work will span production model serving, GPU cluster management, and the infrastructure supporting training and evaluation. You’ll decide how workloads share compute, diagnose performance bottlenecks, and build the deployment and observability tools that help the team move quickly with confidence.
We’re looking for a hands-on engineer who has deployed and operated models in production, understands the demands of GPU workloads, and can take a system from initial design through rollout, incidents, and ongoing improvement. You’ll partner closely with researchers and our platform engineers, with ownership of the systems you build.
What you’ll do
Build and own model-serving infrastructure. Take models from checkpoint to production, with repeatable deployment pipelines, request routing, autoscaling, fallback paths, and controlled rollouts and rollbacks across regions.
Manage GPU clusters and workload scheduling. Improve resource allocation across inference, training, and evaluation. Build scheduling policies around workload priority, quotas, hardware topology, and recovery requirements so online services stay responsive while other workloads make productive use of capacity.
Improve inference performance. Profile real workloads and improve latency, throughput, and memory efficiency. Evaluate batching, KV-cache management, quantization, speculative decoding, and parallelism strategies against production traffic and quality requirements.
Support distributed training and model iteration. Give researchers reliable ways to launch fine-tuning and training jobs, manage model artifacts, save and restore checkpoints, and move validated models into serving. Reduce time lost to failed jobs, slow data loading, and manual setup.
Make deployments observable and incidents traceable. Connect application requests and sessions to the exact model, deployment configuration, and worker that served them. Build dashboards and alerts covering model latency, queueing, errors, GPU health, memory pressure, and workload performance.
Own production reliability. Define service objectives, investigate incidents across the application, inference engine, GPU, and network layers, and build recovery procedures that work. Turn recurring failures into fixes, automated checks, and useful runbooks.
Make compute costs actionable. Track GPU usage, idle capacity, and inference cost by model and workload. Use capacity forecasts and measured performance to guide deployment choices and improve cost per successful request without sacrificing quality or reliability.
-
Build a platform the model team can use independently. Automate provisioning, configuration, benchmarking, and releases. Partner with the engineers behind ASR, note generation, Evidence, and Dictate so new models can be deployed and evaluated through consistent, well-supported workflows.
What you'll need
A strong engineering foundation. Hands-on AI infrastructure experience. At least 1 year building and operating infrastructure for large language models, including model deployment, inference serving, or distributed training. You can design the system, write the code, and own it in production.
Production model deployment experience. You’ve deployed and maintained LLMs or other demanding ML workloads using engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or comparable systems. You understand the work between loading a model and running a reliable service.
GPU cluster and orchestration experience. You’ve managed GPU workloads using Kubernetes, Slurm, or an equivalent platform, with practical experience in scheduling, resource allocation, capacity planning, and failure recovery.
Performance debugging skills. You can use traces, metrics, and profiling tools to distinguish compute, memory, communication, and scheduling bottlenecks. You understand how batch size, context length, precision, and multi-GPU execution affect performance and cost.
Strong software and systems skills. You’re proficient in Python and comfortable with backend or systems development in Go, C++, Rust, or a comparable language. You have practical experience with Linux, containers, deployment automation, and distributed services.
Operational ownership. You’ve owned production incidents, built useful monitoring, and made releases recoverable. You can explain the trade-offs behind a design and work effectively with researchers, product engineers, and infrastructure partners.
Nice to have
Experience with distributed training frameworks such as PyTorch FSDP or Megatron, or infrastructure for reinforcement learning and rollout generation.
Experience tuning inference engines, serving MoE models, or implementing quantization, speculative decoding, and prefill/decode disaggregation.
Familiarity with GPU interconnects, NCCL, RDMA, topology-aware scheduling, or diagnosing multi-node communication problems.
CUDA or Triton kernel development, contributions to AI infrastructure projects, or experience building cluster operators and scheduling integrations.
Experience operating infrastructure across multiple regions or providers, particularly for healthcare or other sensitive production workloads.
How we show up
Build for the next decade, not next quarter. Our targets are outrageous on purpose. The world's health doesn't have the luxury of incrementalism.
Lead, don't wait. We treat tomorrow's problems today. Sometimes we build what's needed before it's wanted, and we're fine with that.
Follow the evidence. Trust the patient. We pursue truth relentlessly. But when the subjective and objective disagree, we treat the patient, not the numbers. Ego is a comorbidity we can't afford.
Own the outcome. Everyone here carries the company. Raise problems with solutions, solve them end-to-end, and never be a bystander.
Ship, measure, go again. A button today, a workflow tomorrow. More iterations beat better planning. We're precise at pace, not reckless.
Live in clinicians' reality. Not the ideal workflow, the twenty-patients-before-lunch actual one. We build for exhausted humans, and we'd better be decent ones while we do it.
Why Heidi?
You’ll join a team focused on real-world impact over imaginary valuations and glossy PR. We live and breathe the challenges of modern health systems, and are laser-focused on exacting the change we’d like to see. We’re medicos, engineers, builders, and designers who’ve felt the moral and practical toll of what non-care feels like. True A-players progress extremely fast here.
The nature of the scale-up game is demanding, but we value sustainable performance and mental health. You're trusted to perform, and you set your schedule. We operate on outcomes > inputs, not process theatre. We all take the bins out, metaphorically and literally.
Building what we’re building isn’t always easy. But we didn’t choose easy, we chose to build something that actually matters. We hold ourselves to a higher standard because healthcare demands it. If you join Heidi, you recognise that the deeper question isn’t whether AI can solve the global healthcare crisis, but whose hands will shape it. The work is hard, but you will trust and admire the people you work beside, and rest easy knowing you’re doing the defining work of your career.
We take care of you.
We offer a $1,000 annual learning and development budget, a $150/month health and wellness allowance, a $500 home office budget, 26 weeks paid primary parental leave and 18 weeks paid secondary parental leave, fertility support up to $10,000, four weeks of work from anywhere per year, and serious equity.




