Adlin-ScienceAdlin-Science

Senior LLMOps, RAG & infrastructure

Paris, Ile-de-France, FranceOn-site or hybridSeniorFULL_TIMEtodayFresh

Location: Paris Type de contrat : Full time permanent position (CDI) Start date: As soon as possible Compensation: Based on experience and profile Adlin Science develops a platform for the governance, quality management and exploitation of multimodal biomedical data. We are strengthening the Data team with a senior pro

LLMsRAGembeddingsDockerCI/CDobservability

Description

Location: Paris Type de contrat : Full-time permanent position (CDI) Start date: As soon as possible Compensation: Based on experience and profile

Adlin Science develops a platform for the governance, quality management and exploitation of multimodal biomedical data. We are strengthening the Data team with a senior profile capable of building the platform’s LLM/IA foundation.

The role is not limited to model serving. It covers usage architecture, selection and qualification of open models, local RAG, evaluation, inference optimization, observability and real-world operation on sensitive data

Role positioning

  • You are the LLM technical referent within the Data team and work closely with the Backend, DevOps, Security, Product and domain expert teams.
  • You design the components specific to the LLM chain and integrate them into the existing Adlin architecture.
  • You do not redefine infrastructure, backend or security standards on your own: you rely on the choices, tools and constraints defined with the responsible teams.
  • You prioritize installable, auditable and maintainable solutions, with no dependency on an external API in production.

Main Responsibilities

  • Select, test and qualify open models compatible with local deployment and licensing, confidentiality and redistribution constraints.
  • Define a modular architecture allowing the model, inference engine or RAG strategy to be changed without excessive coupling.
  • Implement local RAG: embeddings, lexical and vector search, hybrid search, reranking, metadata filters and citation management.
  • Deploy and operate local inference engines such as vLLM, llama.cpp, TGI, TensorRT-LLM, Triton or equivalent solutions depending on the need
  • Optimize latency, throughput, memory consumption and stability through quantization, continuous batching, KV-cache management, parallelism and appropriate format choices.
  • Set up load, resilience and non-regression tests before each production release.
  • Build pipelines for packaging, versioning, validation, promotion and rollback of models, prompts, embeddings, indexes and configurations
  • Define a controlled process for importing models and dependencies into the secure environment: signed artifacts, integrity checks, inventory, vulnerability scans and approval procedure.
  • Containerize components and integrate them with the CI/CD and orchestration tools selected by the Backend / DevOps teams
  • Produce runbooks, incident procedures, architecture files, test evidence and the elements required for security and quality reviews.

Profile

🧰 Required Skills

Technical

  • Hands-on experience deploying and operating LLMs or NLP/ML systems under strong performance constraints.
  • Excellent command of Python, PyTorch, the Transformers ecosystem and the architectural principles of language models.
  • Mastery of at least one high-performance inference engine and experience with quantization and GPU sizing.
  • Practical experience with RAG architectures, embeddings, hybrid search, reranking and evaluation of generative systems.
  • Strong Docker, Linux, CI/CD, observability and artifact management skills in controlled environments.
  • Security culture: secrets, access control, encryption, isolation, software supply chain and sensitive data processing.
  • Ability to document, explain trade-offs and work with multiple teams without creating isolated technical debt. Nice to have
  • Kubernetes, Terraform, MLflow, Prometheus/Grafana, OpenSearch and locally deployable vector databases such as Qdrant, Milvus or pgvector.
  • PEFT/LoRA/QLoRA fine-tuning, distillation, distributed training and preparation of supervised datasets.
  • Knowledge of healthcare sector constraints, GDPR, sensitive data and regulated environments.
  • Experience with air-gapped, on-premise, appliance or edge environments, including controlled import procedures.
  • Open-source contribution or experience reading and critically evaluating scientific papers.

Functional

  • Autonomy and decision-making: you know how to scope your work and move forward even when things are not fully defined.
  • Proactive mindset and critical thinking.
  • Comfortable in dynamic, evolving environments: you thrive when processes are still being shaped.
  • Open and constructive communication: you can break down technical concepts for non-technical audiences, share knowledge freely and foster collaborative dialogue.

🎓 Profile

  • Master’s degree or PhD in computer science, AI, machine learning or a related field, or equivalent experience.
  • Significant experience in ML Engineering, MLOps or AI infrastructure, with at least one LLM deployment actually operated in production.
  • Hands-on profile, able to prototype, code, benchmark, diagnose and industrialize
  • Rigor, autonomy, team spirit and ability to work in a context where traceability and quality come first.
  • Fluent technical English.

Recruitment process

  • Interview with the Data team
  • Interview with Julien, our CTO
  • Interview with David, our CDO

Company

Digital solutions for multi-modal data analysis in scientific and medical research.

Frequently asked

Is this Senior LLMOps, RAG & infrastructure role remote?

This role is based in Paris, Ile-de-France, France.

Related