Research Scientist — Reinforcement Learning (Foundation Models)
ParisRemoteFULL_TIMEtodayFresh
The role We’re hiring a Research Scientist (Reinforcement Learning) to help bring RL into the core of how foundation models are adapted and improved for industrial use. In industry, models don’t live in isolation: domain experts validate, correct, and act on model outputs.
Description
The role
We’re hiring a Research Scientist (Reinforcement Learning) to help bring RL into the core of how foundation models are adapted and improved for industrial use.
In industry, models don’t live in isolation: domain experts validate, correct, and act on model outputs. We want RL to leverage that expert feedback not only as a post‑training patch, but increasingly earlier in the pipeline, shaping objectives, training signals, and adaptation strategies.
What you’ll work on
Human‑in‑the‑loop reinforcement learning
- Turn expert validation/correction into a reliable learning signal.
- Design feedback interfaces/signals that are practical in real operational settings.
RL for industrial foundation models
- Develop RL methods that sit on top of (or integrate with) foundation models used in production.
- Explore ways for RL to intervene earlier in the chain (not just after deployment).
From research to deployment
- Build evaluation protocols aligned with real constraints: robustness, uncertainty reduction, safety, auditability, and cost of error.
- Work closely with scientists/engineers to ship demonstrators that connect benchmarks to field outcomes.
Profile
What we’re looking for
- PhD (preferred) or equivalent research experience in Reinforcement Learning / Machine Learning .
- Strong foundations in RL (e.g., policy optimization, off‑policy learning, offline RL, exploration, credit assignment).
- Ability to design rigorous experiments, debug failure modes, and iterate fast with scientific discipline.
- Strong programming skills (Python; deep learning stack such as PyTorch).
Nice to have
- Experience with real‑world RL constraints (noisy/limited feedback, safety requirements, deployment considerations).
- Comfort with complex data modalities (time series, scientific/industrial signals, multimodal setups).
- Publications or open research artifacts in RL / sequential decision‑making.
Why join
- Work on RL problems that matter in the real world: expert feedback loops, uncertainty reduction, and mission‑critical constraints .
- A research culture that values clarity, rigor, and humility and that connects fundamental ideas to deployable systems.
- High ownership in a small team: you’ll shape direction, not just execute tasks.
Recruitment process
- Recruitment prescreen (30-45min)
- Scientific deep dive (remote-45min)
- Half day of scientific interview (Architecture - Coding - Research talk) + Culture fit
- References call
Company
Foundation models for scientific and industrial data.
Frequently asked
Is this Research Scientist — Reinforcement Learning (Foundation Models) role remote?
Yes — this role is remote-friendly (Paris).