Lesson 4 of 6 · 49 min

Responsible AI & risk

The NIST AI RMF (GOVERN/MAP/MEASURE/MANAGE) and its GenAI profile as the spine, turning risk frameworks into launch artifacts, multi-layer guardrails (NeMo’s five rail types), and red-teaming a probabilistic product — with the convergent “Risk Report” every frontier lab now ships.

Frameworks anchor practice — they don’t replace it

Responsible AI is where PMs either sound mature or sound like they have a checklist. The frameworks (NIST AI RMF, Microsoft’s Responsible AI Standard, Anthropic’s RSP) give you a shared vocabulary and a shape — but NIST deliberately provides no quantitative thresholds, so the concrete go/no-go numbers must come from your own policy. The senior stance: adopt NIST as the spine, write your company-specific gates in the same shape, and link every gate to a verifiable artifact. “The responsible-AI team will review it” is not a gate; a signed go/no-go record with a named owner is.
The NIST AI Risk Management Framework 1.0 (Jan 2023) structures AI risk into four functions, and a PM should anchor every launch artifact to one. GOVERN fixes responsibility and decision rights — without it, everything else floats (no named launcher, no audit chain). MAP contextualises risk to your specific use case — a generic model eval is not a substitute for a context-of-use analysis. MEASURE turns risk hypotheses into measured quantities, with thresholds that block release (subcategory Measure 2.6 requires systems be “regularly evaluated for safety risks”). MANAGE forces an explicit go/no-go and a paper trail — Manage 1.1 is literally the decision of whether deployment proceeds. The whole point: each function must produce a tangible object before release.
code
1NIST AI RMF 1.0 — four functions, one artifact each (what the PM owns)23  FUNCTION   WHAT IT FIXES                LAUNCH ARTIFACT          FAILURE IF SKIPPED4  --------   --------------------------   ----------------------   -------------------5  GOVERN     responsibility / decision    one-page governance      no named launcher,6             rights                       memo (who signs off)     no audit chain7  MAP        context of use               context-of-use para +    generic eval; wrong8                                          risk inventory           risks measured9  MEASURE    risk -> measured quantity    eval suite + release-    "shipped; quality10                                          blocking threshold       TBD" tickets11  MANAGE     explicit go/no-go            signed Manage 1.1        deployment by12                                          go/no-go record          momentum
The companion NIST AI 600-1 Generative AI Profile (July 2024) names twelve unique-or-exacerbated GenAI risks, and it’s the only free, public PM-grade red-team checklist that maps to recognised categories. For each of the twelve — confabulation (verify sources/citations pre-deployment), data privacy, harmful bias/homogenization, human-AI configuration (track anthropomorphization), information integrity, intellectual property, CBRN, dangerous/violent/hateful content, environmental impact, information security, obscene/abusive content, and value-chain integration — write a one-line answer to “what does our eval catch here?” If the answer is “nothing,” that’s a launch blocker. Interview angle. Asked “how do you build responsible AI into the spec instead of auditing for it later,” cite the GenAI profile categories mapped to eval cases in the PRD — that’s building-in; a post-hoc bias audit is auditing-for.
How Microsoft Approaches AI Red Teaming | BRK223Microsoft Developer

Capability thresholds and release plans: turning policy into a gate

Anthropic’s Responsible Scaling Policy is the cleanest example of a concrete gate pattern: capability thresholds map to AI Safety Levels (ASLs), and a given ASL mandates a specific safeguard set before deployment. ASL-3 deployment safeguards are a defense-in-depth stack — tiered access controls, real-time prompt/completion classifiers, asynchronous monitoring classifiers, and post-hoc jailbreak detection with rapid-response patching. The triggering thresholds are explicit (e.g. CBRN uplift to a moderately-resourced state program; the ability to fully automate entry-level AI R&D). The PM lesson: hard thresholds beat vague “high risk” labels — define your own tiers (jailbreak-rate ceilings, hallucination floors, capability eval scores) and refuse promotion without the matching safeguards.
Microsoft’s Responsible AI Standard v2 (June 2022) makes the artifacts concrete with three launch-blocking deliverables: an Impact Assessment (completed early, evaluates impact on people/org/society), a Responsible Release Plan (with specific release criteria), and a Pre-release Evaluation (documented results against six goals: Accountability, Transparency, Fairness, Reliability & Safety, Privacy & Security, Inclusiveness). And Microsoft’s Aether Sensitive Uses review can gate deployments or decline customer engagements outright — its “laser studies” delivered pre-release mitigations for GPT-4 and GitHub Copilot months before GA. The transferable move for any org: a lightweight review board with three powers — approve, approve-with-conditions, halt — that meets weekly.
The strongest cross-company signal in the research is convergence on a Risk Report: Anthropic’s RSP capability thresholds, Microsoft’s Responsible Release Plan, Meta’s “thorough risk assessments and performance testing” for Llama 3.2, and OpenAI’s Preparedness Framework (tracked categories — cyber, CBRN, persuasion, model autonomy — each at low/medium/high/critical) all converge on the same pre-release artifact: a report that names categories, levels, mitigations, and residual uncertainty. Anthropic shipped Claude Opus 4.6 (Feb 2026) under ASL-3 with a public Sabotage Risk Report. The cheapest governance investment a PM can make is a one-page Risk Report attached to every release.

Guardrails: defense in depth, not a single filter

A guardrail is not “a safety filter” — it’s a multi-layer system, and saying “the model is safe” is never sufficient. NVIDIA NeMo Guardrails organises protection into five rail types, each catching a different attack surface: input rails (jailbreak detection, prompt-injection, moderation at the gateway), dialog rails (allowed conversation flows across turns), retrieval rails (filter/validate knowledge-base results), execution rails (gate tool calls — e.g. require human approval for any external action over a threshold), and output rails (fact-check, hallucination detection, PII masking before the response leaves). Guardrails AI adds reusable Validators that turn qualitative criteria into CI-friendly checks. The rule: ship at least input + output rails on every external surface, and dialog + retrieval rails on any multi-turn or knowledge-grounded product.
code
1GUARDRAILS = FIVE RAIL TYPES (NeMo) — ship >=2 layers per path23  RAIL         WHAT IT DOES                          PM USE CASE4  ----------   -----------------------------------   -----------------------------5  Input        jailbreak / prompt-injection / mod.   block "ignore previous6                                                     instructions" at the gateway7  Dialog       allowed multi-turn conversation       refuse off-topic persuasion8               shapes                                requests9  Retrieval    filter / validate KB results          only return cleared-provenance10                                                     docs11  Execution    gate tool calls / actions             human approval for any payment12                                                     or external API > threshold13  Output       fact-check, hallucination, PII mask   strip PII before response ships1415  Design so an on-call PM can DISABLE any single rail without taking the16  feature down. + Microsoft AI Bug Bar: Critical = 15min, Important = 4hr.
Severity matters as much as coverage. Microsoft’s AI Bug Bar classifies AI vulnerabilities into four levels — Critical / Important / Moderate / Low — and the examples are instructive: prompt injection with data exfiltration and no user interaction is Critical; membership inference is Moderate-to-Low depending on training-data sensitivity. Copy this taxonomy and map an on-call response time to each level (Critical = 15 min, Important = 4 hours). The design principle behind both NeMo’s five rails and the four-level bar is the same: no single safety mechanism is sufficient, so you build redundancy at the input, model, output, and post-hoc layers — and you make each layer independently disable-able so an incident doesn’t force you to take the whole feature down.

Red-teaming a probabilistic product

Red-teaming a probabilistic product is not “we tried some bad prompts” — coverage is a function of taxonomy breadth × methodology diversity × write-back into the next eval set, not tester-hours. Anthropic’s “Challenges in red teaming AI systems” lays out the trade-offs: human red teams surface nuanced failures but are slow and small-scale; automated / LLM-elicited red teams scale across known taxonomies but miss failures that need expert context; and evaluation gaps persist — models may “sandbag” (deliberately underperform) or be under-elicited (their true capability is higher than the eval reaches). The two repeated recommendations: combine expert human teams with automated scale, and use multiple, partially-overlapping methods because no single approach covers the surface.
The convergent frontier-lab pattern: pick a public harm taxonomy first, run a multi-method team, write a Risk Report. Meta evaluated fine-tuned Llama 3.2 against privacy, CBRNE, child sexual exploitation, and violent crime, and partnered with MLCommons to adopt its hazard taxonomy; Microsoft ships an Azure AI Red Teaming Agent with a curated seed dataset of risk content. The PM move is to name the categories you covered and the ones you did not — “we red-teamed for safety” is meaningless without the taxonomy. Interview angle. “How would you design guardrails for an agentic system that can take actions on a user’s behalf?” → least-privilege scoped permissions per action, human confirmation for irreversible actions (payments, emails), spending caps and per-session rate limits with numbers, and a documented kill switch + on-call rotation. Listing rails without a kill switch is the weak answer.
Finally, red-teaming feeds the eval set: every Critical finding becomes a permanent negative test case, so the same failure can never silently regress back in. This closes the loop with L1 and L3 — the eval portfolio is where responsible-AI findings live permanently. A pre-red-team checklist a PM can run: (1) select a public taxonomy (MLCommons, NIST AI 600-1), (2) commission an expert human team for the highest-impact categories, (3) run automated LLM-driven fuzzing at scale, (4) probe for under-elicitation, (5) cross-validate top findings between methods, (6) route Critical findings to a remediation owner with a deadline, (7) roll all findings into the Risk Report attached to release.
paperNIST AI 600-1 — Generative AI Profile (the 12 risk categories)NISTarticleAnthropic — Challenges in red teaming AI systemsAnthropicrepoNVIDIA NeMo Guardrails — five rail types in the Colang DSLNVIDIAdocsMicrosoft — Vulnerability Severity Classification for AI Systems (AI Bug Bar)Microsoft MSRC

Checkpoint

A PM says “we’ll comply with the NIST AI RMF, so we’re covered on go/no-go thresholds.” What’s the issue?

ANIST RMF is outdated and superseded by the EU AI ActBNIST provides the shape (GOVERN/MAP/MEASURE/MANAGE) but no quantitative thresholds — the concrete go/no-go numbers must come from your own policy, linked to verifiable artifactsCNIST RMF only applies to government systems, not products
Sign up free to answer and see why

Checkpoint

Your team ships an output content-moderation filter and calls the launch “safe.” A security reviewer pushes back. Why?

AOne output filter is a single rail — it can’t stop prompt injection on the input, poisoned retrieval, or unsafe tool calls; you need defense in depth (≥2 rails per path) and a severity-tiered responseBOutput filters add too much latency to be worth itCContent moderation should be done by humans, not a filter
Sign up free to answer and see why

Checkpoint

Asked to design guardrails for an agent that can send emails and make purchases on a user’s behalf, which answer is strongest?

AAdd safety filters and monitor for misuseBLeast-privilege scoped permissions per action, human confirmation for irreversible actions, spending caps and per-session rate limits with numbers, and a documented kill switch + on-call rotationCOnly allow the agent to take “safe” actions
Sign up free to answer and see why

Checkpoint

A PM reports “we red-teamed the model for safety before launch.” What does a senior reviewer ask to know if that’s meaningful?

AHow many hours the red team spentBHow many people were on the red teamCWhich public harm taxonomy was used, which methods (human + automated), what was covered AND not covered, and whether findings became permanent eval cases
Sign up free to answer and see why

Checkpoint

Across Anthropic, Microsoft, Meta, and OpenAI, what single pre-release artifact do their responsible-AI processes converge on?

AA public leaderboard scoreBA Risk Report naming categories, risk levels, mitigations, and residual uncertainty (plus rollback) — attached to the releaseCA marketing launch blog post
Sign up free to answer and see why

Interview prep

Responsible-AI questions test two literacies at once: ethics (can you hold a launch when a subgroup is harmed) and architecture (can you build constraint into the design, not the audit). Product School frames it as a maturity signal — strong candidates treat responsible AI as “the line between trusted products and unpredictable ones,” weak ones say “we have a checklist.” Pair every answer with an affected-group segmentation, a quantified harm, a launch gate, and a named owner.
  1. 01“How do you build fairness/privacy/transparency into the spec, not audit later?” → map NIST AI 600-1 risk categories to eval cases in the PRD, with a launch-gate field and an owner.
  2. 02“Model works for 90% but fails for one demographic — what do you do, on what timeline?” → stratified eval, a named harm-rate threshold, shadow/human-review that slice, and a remediation owner with a deadline.
  3. 03“Design guardrails for an agentic system.” → scoped least-privilege permissions, human confirm for irreversible actions, spend caps + rate limits (with numbers), documented kill switch + on-call.
  4. 04“What framework anchors your launch?” → NIST AI RMF (GOVERN/MAP/MEASURE/MANAGE) as the spine; company-specific thresholds in the same shape, linked to artifacts.
  5. 05“What goes in a pre-release Risk Report?” → categories (NIST 600-1) × risk levels × mitigations × residual uncertainty × rollback plan, one page, named owners.
  6. 06“How do you red-team a probabilistic product?” → pick a public taxonomy, combine human + automated methods, name what you didn’t cover, fold Critical findings into the eval set.
  7. 07“How do you triage an AI vulnerability?” → a four-level bug bar (Critical/Important/Moderate/Low) with response SLAs; prompt-injection-with-exfiltration is Critical.
  8. 08“Give a real example where an ethics concern changed a product decision.” → name the harm, the affected group, the gate you set, and the launch delay you accepted.
Going deeper: “who has final authority to delay launch on an ethics ground?” (a named owner / review board with halt power — Microsoft’s Aether model); “if legal says it’s fine but the advocacy team says it’s harmful, what do you do?” (escalate to the board, document the residual risk in the Risk Report, and make the call explicit rather than silent); and “how do you detect the harmed 10% without flagging it for discrimination?” (per-slice quality monitoring with privacy-aware aggregation). The thread: responsible AI is a launch gate with an owner and an artifact, not a sentiment.

Could you anchor a launch to the NIST functions, design multi-layer guardrails with a kill switch, and outline a multi-method red-team plus a one-page Risk Report?

New to itGetting thereConfident

Takeaways

  • NIST AI RMF (GOVERN/MAP/MEASURE/MANAGE) is the spine — but it gives shape, not numbers; your policy sets the thresholds.
  • Map the 12 NIST GenAI-profile risks to eval cases in the PRD — that’s building responsible AI in, not auditing for it.
  • Guardrails are five rail types (input/dialog/retrieval/execution/output) — ship ≥2 per path, each independently disable-able.
  • Triage with a four-level bug bar (Critical=15min … ); prompt-injection-with-exfiltration is Critical.
  • Red-team coverage = taxonomy breadth × method diversity × write-back; name what you did NOT cover.
  • Every frontier lab converges on a one-page Risk Report: categories × levels × mitigations × residual uncertainty × rollback.

Next: launch gates & incident handling — go/no-go, rollback rehearsal, AI incident response, and post-launch monitoring.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.