Sigma NovaSigma Nova

Research Scientist — Reinforcement Learning (Foundation Models)

ParisRemoteFULL_TIMEtodayFresh

The role We’re hiring a Research Scientist (Reinforcement Learning) to help bring RL into the core of how foundation models are adapted and improved for industrial use. In industry, models don’t live in isolation: domain experts validate, correct, and act on model outputs.

PythonPyTorch

Description

The role

We’re hiring a Research Scientist (Reinforcement Learning) to help bring RL into the core of how foundation models are adapted and improved for industrial use.

In industry, models don’t live in isolation: domain experts validate, correct, and act on model outputs. We want RL to leverage that expert feedback not only as a post‑training patch, but increasingly earlier in the pipeline, shaping objectives, training signals, and adaptation strategies.

What you’ll work on

Human‑in‑the‑loop reinforcement learning

  • Turn expert validation/correction into a reliable learning signal.
  • Design feedback interfaces/signals that are practical in real operational settings.

RL for industrial foundation models

  • Develop RL methods that sit on top of (or integrate with) foundation models used in production.
  • Explore ways for RL to intervene earlier in the chain (not just after deployment).

From research to deployment

  • Build evaluation protocols aligned with real constraints: robustness, uncertainty reduction, safety, auditability, and cost of error.
  • Work closely with scientists/engineers to ship demonstrators that connect benchmarks to field outcomes.

Profile

What we’re looking for

  • PhD (preferred) or equivalent research experience in Reinforcement Learning / Machine Learning .
  • Strong foundations in RL (e.g., policy optimization, off‑policy learning, offline RL, exploration, credit assignment).
  • Ability to design rigorous experiments, debug failure modes, and iterate fast with scientific discipline.
  • Strong programming skills (Python; deep learning stack such as PyTorch).

Nice to have

  • Experience with real‑world RL constraints (noisy/limited feedback, safety requirements, deployment considerations).
  • Comfort with complex data modalities (time series, scientific/industrial signals, multimodal setups).
  • Publications or open research artifacts in RL / sequential decision‑making.

Why join

  • Work on RL problems that matter in the real world: expert feedback loops, uncertainty reduction, and mission‑critical constraints .
  • A research culture that values clarity, rigor, and humility and that connects fundamental ideas to deployable systems.
  • High ownership in a small team: you’ll shape direction, not just execute tasks.

Recruitment process

  • Recruitment prescreen (30-45min)
  • Scientific deep dive (remote-45min)
  • Half day of scientific interview (Architecture - Coding - Research talk) + Culture fit
  • References call

Company

Foundation models for scientific and industrial data.

Frequently asked

Is this Research Scientist — Reinforcement Learning (Foundation Models) role remote?

Yes — this role is remote-friendly (Paris).

Related