ML Research Engineer - Post-training (LLMs)
BengaluruRemoteMidFULL_TIMEtoday
ML Research Engineer; Post training (LLMs) Bengaluru · Full time · Experience: 4–8 yrs India's healthcare runs in twenty two languages, on handwritten prescriptions and ten minute consults, and the models that should serve it are trained on the English internet. We're fixing that, in the open.
ML Research Engineer; Post-training (LLMs)
Bengaluru · Full-time · Experience: 4–8 yrs
India's healthcare runs in twenty-two languages, on handwritten prescriptions and ten-minute consults, and the models that should serve it are trained on the English internet. We're fixing that, in the open.
About EkaCare
EkaCare is India's connected healthcare platform: an EMR that doctors run their practices on, a personal health record used by millions of Indians, and one of the deepest integrations with India's ABDM digital-health rails. Our Parrotlet family of medical models already serves Indian doctors in production, and we open-source our work where it counts.
The role
Post-training is where a base model becomes a doctor's tool, and where most medical models quietly fail. You'll turn a strong pre-trained base into a model that follows instructions, knows what it doesn't know, stays safe in a clinical setting, and does it in a dozen Indian languages.
What you'll do
- Design SFT/IFT data mixes and chat templates for clinical tasks, and reformat the world's medical data to match.
- Run preference optimisation (DPO/ORPO/GRPO-class) and reward-model training; own the ablation grid.
- Build RL loops with verifiable medical rewards, with practising physicians in the loop. Real doctor-in-the-loop, not proxy labels.
- Create agentic and tool-use training data for healthcare workflows.
- Live in the eval → error-analysis → iterate loop; red-team your own model before the world does.
What we look for
- 2–4 years in ML with hands-on post-training of ≥7B open-weights models; you've shipped SFT plus at least one preference-optimisation method end to end, not a notebook demo.
- Fluency with the open post-training stack (TRL / NeMo RL / OpenRLHF; vLLM for rollouts).
- Strong empirical taste: ablation discipline, LLM-judge literacy, contamination paranoia.
- Solid engineering, Python, distributed-training basics, comfort in a fast codebase.
Bonus
- PPO/GRPO at scale; reward-hacking war stories.
- Multilingual or medical/clinical alignment work.
- Open-source contributions people actually use.
Frequently asked
Is this ML Research Engineer - Post-training (LLMs) role remote?
Yes — this role is remote-friendly (Bengaluru).