AI Assurance Testing and Governance Framework
Aligned to NIST AI RMF 1.0 (NIST AI 100-1) and NIST AI 600-1 (Generative AI Profile
Date: January 13, 2026
1. Executive Summary
Framework Overview
This AI Governance Framework provides a comprehensive, risk-based approach to developing, deploying, and monitoring artificial intelligence systems. It operationalizes the NIST AI Risk Management Framework (AI RMF 1.0) and incorporates specialized controls from NIST AI 600-1 for Generative AI systems.
Framework Statistics:
32 Core Controls organized into 15 categories
3-Tier Risk Classification system (Low/Medium/High)
4 Lifecycle Phases aligned to NIST AI RMF (GOVERN-MAP-MEASURE-MANAGE)
10 Comprehensive Appendices including templates, playbooks, and training curriculum
Global Compliance Coverage including EU AI Act, GDPR, CCPA, ISO 42001
Why This Framework Matters
Business Value:
Accelerates time-to-market with clear approval pathways
Reduces regulatory risk through comprehensive compliance mapping
Protects brand reputation with proactive risk management
Enables innovation within defined guardrails
Risk Management:
Systematic identification and mitigation of AI-specific risks
Proportional controls based on risk level (avoiding one-size-fits-all)
Continuous monitoring and improvement mechanisms
Incident response preparedness
Regulatory Compliance:
Aligned with EU AI Act high-risk system requirements
Maps to GDPR, CCPA, and sector-specific regulations
Supports SOC 2, ISO 27001, and ISO 42001 certification
Demonstrates due diligence for liability protection
Framework Philosophy
Risk-Based, Not Checklist-Based: Controls scale to the actual risk level of each AI system. A Tier 1 internal summarizer requires minimal governance; a Tier 3 public-facing system requires extensive validation.
Outcome-Focused: Teams can implement controls differently by context, as long as they achieve required outcomes and provide evidence.
Lifecycle-Oriented: Every AI system progresses through GOVERN → MAP → MEASURE → MANAGE with clear gates and artifacts at each stage.
GenAI-Aware: Specialized controls for hallucination, prompt injection, provenance, and transparency address unique generative AI risks.
Who Should Use This Framework
Required Users:
AI/ML Engineers and Data Scientists
Product Managers deploying AI features
Business owners of AI systems
Vendors/third-parties providing AI services
Governance and Oversight:
AI Risk Leads
AI Review Board (ARB) members
Security, Privacy, and Legal reviewers
Domain experts and quality assurance
Executive risk owners
Support Functions:
Compliance and audit teams
Procurement (vendor assessment)
Training and enablement
Incident response teams
Quick Start Guide
For AI System Owners:
Complete AI Intake Form (Section 9.1)
Calculate risk tier using scoring rubric (Section 6)
Prepare required artifacts for your tier (Section 7)
Submit for review and approval (Section 8)
Deploy with monitoring and post-launch review
For Reviewers:
Review assigned artifacts for completeness
Validate risk tier assignment
Assess controls against tier requirements
Provide approval or request remediation
Schedule post-launch review
For Executives:
Understand risk tiering and automatic overrides (Section 6)
Review Tier 3 systems requiring executive risk acceptance
Monitor program metrics and incident trends
Ensure adequate resources for governance program
Recommended Framework Maintenance
Organizations implementing this framework should maintain it as a living document with the following recommended schedule:
Annual comprehensive review (January recommended)
Regulatory change response (within 90 days of major changes)
Incident driven updates (within 30 days of incidents revealing gaps)
Version control with full change tracking
Success Metrics
Organizations implementing this framework should track:
Coverage: % of AI systems with completed governance artifacts
Velocity: Average time from intake to approval by tier
Quality: % of systems passing TEVV on first attempt
Safety: Number and severity of post-launch incidents
Compliance: Audit findings and regulatory observations
2. How to Use This Document
Document Structure
Sections 1-8: Core Governance Standard
Publish these sections as your official AI governance policy
Required reading for all AI system stakeholders
Defines principles, roles, risk tiers, and workflow
Section 9: Templates
Use these templates for every AI system
Copy/paste and complete as required by tier
Store completed templates in governance system
Appendix A: Examples
Real-world examples of completed templates
Demonstrates what "good" looks like at each tier
Reference when creating your own artifacts
Appendices B-J: Supporting Materials
Reference materials, tools, and guidance
Use as needed based on your role and system type
Updated independently of core standard
Navigation Guide
If you are...
Deploying your first AI system: → Start with Section 3 (Scope), then Section 9.1 (Intake Form), then Section 6 (Tiering)
An AI Risk Lead: → Master Sections 6-8 (Tiering, Artifacts, Workflow), use Appendix D (Risk Assessment Template)
A Security Reviewer: → Focus on Section 7 (Required Artifacts - Threat Model), use Appendix F (Incident Response)
An Executive: → Read Section 4 (Principles), Section 6.1 (Tier 3 Overrides), review Appendix A.3 (Tier 3 Example)
Assessing a Vendor: → Use Appendix H (Vendor Assessment Questionnaire)
Responding to an Incident: → Jump to Appendix F (Incident Response Playbook)
Ensuring EU AI Act Compliance: → Review Section 12 (Regulatory Landscape) and Appendix I (Prohibited AI)
Templates Are Available In:
Copy/paste format: Sections 9.1-9.6
Fillable forms: In governance system
Spreadsheet format: For complex tracking (e.g., control catalog)
Customization Guidance
Organizations should customize this framework by:
Required Customizations:
Replace
[Organization Name]with your organizationUpdate contact information (Section: Contact Information)
Add your approval authorities and names (Section 5)
Configure governance system URLs and tools
Add organization-specific prohibited use cases
Recommended Customizations:
Adjust tier score thresholds based on risk appetite (Section 6.1)
Add industry-specific controls (e.g., healthcare, financial services)
Customize training curriculum to organizational tools (Appendix G)
Add examples from your organization (Appendix A)
Incorporate existing policies and standards
Prohibited Modifications:
Do not reduce NIST AI RMF alignment
Do not weaken regulatory compliance requirements
Do not remove mandatory controls for Tier 2-3 systems
Do not bypass approval gates without formal exception process
Getting Help
Questions about:
Tiering: Contact AI Risk Lead
Artifacts: Reference Appendix A examples
Compliance: Consult Privacy/Legal team
Technical Implementation: Engage Security team
Training: Contact AI Governance Team
Resources:
Internal Portal: [URL to be configured]
Slack Channel: #ai-governance
Email: ai-governance@[organization].com
Monthly Office Hours: [Calendar link]
3. Purpose
This standard defines how the organization approves, deploys, and monitors AI systems using a risk-based approach. It operationalizes the NIST AI Risk Management Framework (AI RMF) lifecycle “GOVERN, MAP, MEASURE, MANAGE” and adds GenAI-specific controls from NIST AI 600-1 where applicable.
Primary Objectives:
Protect People: Ensure AI systems do not cause unacceptable harm to safety, rights, privacy, or dignity
Enable Innovation: Provide clear pathways for responsible AI deployment at speed
Manage Risk: Identify, assess, and mitigate AI-specific risks proportional to impact
Ensure Compliance: Meet regulatory requirements across jurisdictions (EU AI Act, GDPR, CCPA, etc.)
Build Trust: Demonstrate responsible AI practices to users, regulators, and stakeholders
Drive Accountability: Establish clear ownership and decision rights for AI systems
Relationship to Other Policies:
This AI Governance Standard supplements and works alongside:
Information Security Policy
Data Privacy Policy
Software Development Lifecycle (SDLC) Policy
Vendor Management Policy
Incident Response Policy
Business Continuity Policy
Where conflicts arise, this AI-specific standard takes precedence for AI systems.
4. Scope
In Scope
This standard applies to:
AI System Types:
Machine Learning (ML) systems (classification, regression, clustering, etc.)
Large Language Models (LLMs) and Generative AI (GenAI) systems
Computer vision and image recognition systems
Natural language processing (NLP) systems
Recommendation engines
Predictive analytics and forecasting systems
Autonomous or semi-autonomous decision systems
Any system using statistical models to generate outputs influencing decisions, actions, or user experiences
Deployment Modes:
Internal-use systems (employee-facing tools)
Partner/B2B systems (restricted external access)
Customer/public-facing systems (open external access)
Embedded AI in products or services
API-based AI services
Acquisition Models:
Internally developed AI systems
Vendor-provided AI via API/SaaS (including "copilots," assistants, embedded AI)
Open-source models deployed by organization
Hybrid systems (vendor model + internal fine-tuning)
Lifecycle Stages:
Pre-deployment (development, testing, validation)
Deployment (production launch, scaling)
Operations (monitoring, maintenance, updates)
Decommissioning (retirement, data deletion)
Out of Scope
This standard does NOT apply to:
Simple rule-based systems without machine learning
Statistical reporting and business intelligence (no predictive/generative component)
Traditional software with no AI/ML components
Research and experimentation (until moving to production deployment)
- Exception: Research involving human subjects must follow research ethics protocols
Third-party AI used by individual employees for personal productivity (covered by Acceptable Use Policy)
- Exception: If personal use involves confidential data, additional policies apply
Scope Boundaries
When Does Governance Start?
Development: When moving from research/POC to production-intended development
Vendor: When selecting or onboarding a new AI vendor/service
Updates: When making material changes to existing systems (see Section 8.8)
When Does Governance End?
Upon formal decommissioning and data deletion
Or upon transfer of ownership (with governance transfer documentation)
5. Definitions
Core Definitions
AI System A system that uses machine learning, deep learning, or statistical models to generate outputs that influence decisions, actions, or user experiences. This includes systems that classify, predict, generate, recommend, or assist in decision-making.
Generative AI (GenAI) System An AI system designed to generate new content (text, images, audio, video, code, or other media) in response to prompts or inputs. Includes Large Language Models (LLMs), image generators, code generators, and multimodal systems.
Large Language Model (LLM) A type of GenAI system trained on vast text datasets to understand and generate human-like text. Can perform tasks like question-answering, summarization, translation, code generation, and conversation.
Machine Learning (ML) A subset of AI where systems learn patterns from data to make predictions or decisions without explicit programming for every scenario.
Foundation Model A large-scale AI model trained on broad data that can be adapted to many downstream tasks (e.g., GPT-4, Claude, BERT, DALL-E). Often used as the base for fine-tuning or prompt engineering.
Governance Terms
TEVV (Testing, Evaluation, Validation, and Verification) Comprehensive testing performed before deployment and after material changes to ensure the AI system meets safety, performance, fairness, and security requirements.
Risk Tier Classification of an AI system into Tier 1 (Low), Tier 2 (Medium), or Tier 3 (High) based on potential impact, using a standardized scoring rubric (Section 6).
System Card A structured document describing an AI system's purpose, architecture, data handling, limitations, monitoring, and controls (Template: Section 9.2).
Threat Model Analysis of potential security threats, attack vectors, and mitigations specific to an AI system (Template: Section 9.4).
AI Review Board (ARB) Cross-functional governance body that reviews and approves Tier 2-3 AI systems before deployment.
Executive Risk Owner Senior executive who accepts residual risk for Tier 3 systems after all mitigations are applied.
AI Risk and Trustworthiness Terms
Hallucination When a GenAI system generates false, fabricated, or nonsensical information presented as fact. Particularly concerning when users rely on outputs for decisions.
Prompt Injection An attack where malicious input (prompts) manipulates an AI system to bypass safety controls, reveal confidential information, or perform unintended actions.
Bias Systematic and unfair discrimination against certain groups or individuals in AI system outputs, often stemming from training data or model design.
Fairness The principle that AI systems should treat all individuals and groups equitably, without unjust discrimination based on protected characteristics.
Explainability The degree to which an AI system's decisions or outputs can be understood and interpreted by humans.
Provenance The ability to trace the origin and creation process of AI-generated content, including model version, inputs, and any human edits.
Drift Performance degradation over time as real-world data diverges from training data (data drift) or as model behavior changes (model drift).
Technical Terms
Fine-Tuning Further training a pre-trained model on specific data to adapt it to particular tasks or domains.
Retrieval-Augmented Generation (RAG) A technique where a GenAI system retrieves relevant information from external sources before generating a response, improving accuracy and grounding.
Prompt Engineering The practice of designing effective instructions (prompts) to guide AI system behavior and outputs.
Few-Shot Learning Providing a small number of examples in the prompt to guide the AI system's behavior without formal retraining.
Temperature A parameter controlling randomness in AI-generated outputs (lower = more deterministic, higher = more creative/random).
Token The basic unit of text processing in LLMs (roughly equivalent to words or subwords).
Context Window The maximum amount of text (in tokens) an LLM can process at once, including both input and output.
Regulatory and Standards Terms
NIST AI RMF (Risk Management Framework) US National Institute of Standards and Technology's framework for managing AI risks, organized around four functions: GOVERN, MAP, MEASURE, MANAGE.
EU AI Act European Union regulation (2024) classifying AI systems by risk level and imposing requirements, especially for high-risk systems.
GDPR (General Data Protection Regulation) EU regulation protecting personal data and privacy, with specific provisions affecting AI systems.
CCPA/CPRA (California Consumer Privacy Act / California Privacy Rights Act) California laws providing privacy rights and affecting AI systems that process California residents' personal information.
ISO 42001 International standard for AI Management Systems, providing requirements for responsible development and use of AI.
High-Risk AI System (EU AI Act Definition) AI systems in specified domains (employment, law enforcement, critical infrastructure, etc.) subject to strict EU AI Act requirements.
6. Key Principles
6.1 Risk-Based, Rights-Preserving Governance
Principle: Apply controls proportional to potential impact. Do not deploy systems with unacceptable risks to people's safety, rights, or dignity.
What This Means:
A simple internal summarizer (Tier 1) requires basic documentation and testing
A customer-facing recommendation system (Tier 2) requires comprehensive testing and monitoring
A public-facing eligibility advisor (Tier 3) requires extensive validation, red-team testing, and executive oversight
Non-Negotiable Limits:
Never deploy systems on the EU AI Act prohibited list
Never deploy systems with known, unmitigated risks to safety or fundamental rights
Never proceed without required approvals for the tier
Compliance Evidence:
Documented risk tier score with justification
Required artifacts produced and reviewed for the tier
Documented risk acceptance when residual risk remains (Tier 3)
Post-launch review confirming risk profile remains accurate
6.2 Lifecycle Governance: GOVERN–MAP–MEASURE–MANAGE
Principle: Every AI system must produce artifacts aligned to each NIST AI RMF lifecycle function and pass approval gates before launch.
GOVERN: Establish policies, roles, and risk appetite
Risk tiering and classification
Role assignment and accountability
Policy compliance verification
MAP: Understand context, purpose, and risks
System Card documenting intended use and architecture
Data mapping and privacy assessment
Threat modeling and risk identification
MEASURE: Test and validate before deployment
TEVV testing against defined thresholds
Fairness and bias evaluation
Security and adversarial testing
Performance benchmarking
MANAGE: Monitor, maintain, and respond
Production monitoring and alerting
Incident response and rollback capability
Model updates and retraining
Post-launch reviews and continuous improvement
Compliance Evidence:
System Card (MAP)
Threat Model (MAP)
Risk Analysis (GOVERN/MANAGE)
TEVV Report (MEASURE)
Monitoring/Rollback Plan (MANAGE)
Post-launch review documentation (MANAGE)
6.3 Trustworthy AI Objectives Are Explicit and Measurable
Principle: Define measurable targets for applicable trustworthiness characteristics and document trade-offs.
Eight Trustworthiness Characteristics (NIST AI RMF):
Reliability: System performs consistently and predictably
- Metric examples: Uptime %, error rate, task success rate
Safety: System does not pose unacceptable risks to health, safety, or rights
- Metric examples: Harmful output rate, safety filter effectiveness
Security: System is protected against threats and attacks
- Metric examples: Vulnerability count, prompt injection resistance
Accountability: Clear ownership and responsibility
- Metric examples: Audit trail completeness, incident response time
Transparency: Users understand when and how AI is used
- Metric examples: Disclosure rate, documentation completeness
Explainability: Decisions can be understood and interpreted
- Metric examples: Explanation availability, user comprehension scores
Privacy: Personal data is protected
- Metric examples: PII leakage rate, data minimization compliance
Fairness / Bias Management: Equitable treatment across groups
- Metric examples: Performance parity across demographics, bias audit results
Trade-Off Documentation Required:
Which characteristics are prioritized and why
Documented limitations where perfect performance is unattainable
Accepted trade-offs (e.g., accuracy vs. explainability, safety vs. helpfulness)
Compliance Evidence:
TEVV metrics and thresholds for each applicable characteristic
Documented limitations and trade-offs in System Card
Post-launch monitoring SLOs and alert thresholds
Regular fairness and bias audits (Tier 2-3)
6.4 Outcome-Focused Governance (Not One-Size Checklist)
Principle: Teams may implement controls differently by context, as long as required outcomes and evidence are met.
What This Means:
A team can choose their testing methodology, as long as they meet TEVV thresholds
Different architectures (vendor LLM vs. self-hosted) require different but equivalent controls
Domain-specific requirements (healthcare vs. marketing) drive different implementations
Required Outcomes (Non-Negotiable):
Risk tier accurately reflects actual risk
All required artifacts exist and are complete
Testing meets or exceeds tier thresholds
Monitoring is in place before launch
Approvals are obtained from required roles
Alternative Control Implementation:
Must be documented in System Card or Risk Analysis
Must achieve equivalent or better risk mitigation
Must be approved by relevant reviewers (Security, Privacy, Domain Lead)
Examples: Different encryption methods, alternative testing frameworks, compensating controls
Compliance Evidence:
Outcome-to-evidence mapping documented
Alternative implementations explicitly approved
Regular validation that alternative controls remain effective
6.5 GenAI-Specific Governance (NIST AI 600-1)
Principle: GenAI systems must include additional controls for content provenance and transparency, context-appropriate pre-deployment testing, and disclosure readiness for significant issues.
Why GenAI Is Different:
Hallucination Risk: Can generate convincing but false information
Prompt Injection: Vulnerable to input-based attacks
Data Leakage: May memorize and regurgitate training data
Provenance Challenges: Difficult to track AI-generated content
Rapid Evolution: Models update frequently with changing behavior
Multi-Modal Complexity: Text, image, code generation creates diverse risk surfaces
Additional GenAI Requirements:
For All GenAI Systems:
Content safety filters and refusal mechanisms
Testing for hallucination and grounding
Prompt injection security testing
User disclosure of AI-generated content
For Tier 3 GenAI (Required):
Provenance & Transparency Plan documenting content origin tracking
Adversarial/Red-Team Testing with comprehensive attack scenarios
Disclosure Readiness plan for significant safety or security issues
Source attribution and citation mechanisms (where applicable)
Enhanced monitoring for model update regressions
For Tier 2 GenAI (Recommended):
Simplified provenance tracking
Basic red-team testing scenarios
Content watermarking or metadata (where feasible)
Compliance Evidence:
Provenance & Transparency Plan (Tier 3 GenAI; recommended others)
Adversarial/red-team testing in TEVV (Tier 3; recommended Tier 2)
Disclosure readiness section in Risk Analysis (Tier 3)
GenAI-specific monitoring metrics (hallucination rate, refusal rate, etc.)
7. Roles and Decision Rights
7.1 Minimum Required Roles
Every AI system must have these roles assigned:
Business Owner
Accountability: Intended use, user impact, business value, and outcomes
Responsibilities:
Define business requirements and success criteria
Approve System Card and intended use
Make risk acceptance decisions (with Executive for Tier 3)
Authorize deployment and changes
Own post-launch review and continuous improvement
Required For: All tiers
Engineering Owner
Accountability: Implementation, deployment, monitoring, and technical operations
Responsibilities:
Develop or integrate AI system
Implement security and privacy controls
Create TEVV plan and execute testing
Deploy and configure monitoring
Maintain rollback capability
Respond to technical incidents
Required For: All tiers
AI Risk Lead
Accountability: Governance process, risk assessment, artifact completeness
Responsibilities:
Run intake and tiering process
Validate risk tier assignment
Ensure required artifacts are complete
Coordinate ARB reviews
Maintain governance documentation
Track open risks and mitigations
Report governance metrics
Required For: All tiers (organizational role, not per-system)
Security Reviewer
Accountability: Threat modeling, security controls, abuse resistance
Responsibilities:
Conduct or review threat model
Assess security testing adequacy (prompt injection, access controls, etc.)
Review access control implementation
Approve security aspects before launch
Participate in security incident response
Required For: Tier 2-3
Privacy/Legal Reviewer
Accountability: Data handling compliance, privacy controls, regulatory alignment
Responsibilities:
Review data processing and retention
Assess GDPR, CCPA, or other privacy compliance
Review vendor contracts and DPAs
Approve privacy aspects before launch
Assess regulatory notification requirements
Required For: Tier 2-3
Domain Lead
Accountability: Domain correctness, quality standards, user impact assessment
Responsibilities:
Define domain-specific quality requirements
Review TEVV test coverage and thresholds
Assess output quality and correctness
Validate domain-appropriate use cases
Approve domain aspects before launch
Required For: Tier 2-3
Examples: Clinical lead (healthcare AI), Financial analyst (FinTech AI), HR lead (people analytics)
7.2 Governance Bodies
AI Review Board (ARB)
Composition: Cross-functional representatives (Engineering, Product, Security, Privacy/Legal, Domain, Risk)
Chair: Senior technical or risk leader
Accountability: System approval for Tier 2-3
Responsibilities:
Review risk tier and required evidence
Assess completeness and quality of artifacts
Make launch/no-launch decisions
Approve risk treatment plans
Escalate to Executive Risk Owner (Tier 3)
Track program-level metrics and trends
Meeting Cadence: Weekly or as-needed for Tier 2; within 48hrs for urgent Tier 3
Decision Authority: Consensus preferred; Chair breaks ties
Delegation: ARB may delegate Tier 2 approvals to qualified reviewers
Executive Risk Owner
Accountability: Residual risk acceptance for Tier 3 systems
Who: CTO, CISO, CPO, General Counsel, or equivalent C-level executive
Responsibilities:
Review Tier 3 system risk profile
Accept residual risks after all mitigations
Sign off on deployment authorization
Own escalation for major incidents
Report to Board on high-risk AI deployments
Required For: Tier 3 only
7.3 RACI Matrix
| Activity | Business Owner | Engineering Owner | AI Risk Lead | Security | Privacy/Legal | Domain Lead | ARB | Executive |
| Define Use Case | A | C | I | I | I | C | I | I |
| Risk Tiering | C | C | A | C | C | C | I | I |
| System Card | A | R | C | I | I | C | I | I |
| Threat Model | C | R | C | A | I | I | I | I |
| TEVV Execution | C | A/R | C | C | I | A | I | I |
| Risk Analysis | A | C | R | A | A | A | C | I |
| Tier 1 Approval | A | A | C | I | I | I | I | I |
| Tier 2 Approval | C | C | C | A | A | A | A | I |
| Tier 3 Approval | C | C | C | A | A | A | A | A |
| Deployment | C | A/R | I | C | I | I | I | I |
| Monitoring | C | A/R | C | C | I | C | I | I |
| Incident Response | C | A/R | C | A | A | C | I | A (P1) |
Legend:
R = Responsible (does the work)
A = Accountable (final approval/decision)
C = Consulted (provides input)
I = Informed (kept updated)
7.4 Role Assignment Requirements
Tier 1:
Minimum: Business Owner + Engineering Owner
AI Risk Lead coordinates intake/tiering
Tier 2:
- All Tier 1 roles PLUS: Security, Privacy/Legal, Domain Lead, ARB
Tier 3:
All Tier 2 roles PLUS: Executive Risk Owner
Enhanced ARB (full board, not delegate)
Vendor/Third-Party AI:
Business Owner responsible for vendor relationship
Engineering Owner responsible for integration
Privacy/Legal must review contracts and DPAs
Vendor Assessment (Appendix H) required for new vendors
8. Risk Tiering
8.1 Risk Tiering Rubric
The tiering rubric converts qualitative risk into a repeatable score (0–18). Score each dimension 0–3 and record a short justification.
Dimension D1: External Exposure
Question: Who sees the outputs?
| Score | Description | Examples |
| 0 | Internal team only; restricted access | Dev team testing tool, internal analytics dashboard |
| 1 | Organization-wide internal use | Company-wide chatbot, HR tool, sales assistant |
| 2 | Partners or controlled external parties | B2B API, partner portal, supplier interface |
| 3 | Public-facing or uncontrolled distribution | Public website feature, consumer app, open API |
Dimension D2: Decision / Harm Impact
Question: What happens if it's wrong?
| Score | Description | Examples |
| 0 | Convenience only; no material impact | Meeting summarizer, draft email writer, idea brainstormer |
| 1 | Productivity/efficiency impact; easily reversible | Document classifier, content moderator assist, search ranking |
| 2 | Financial/reputational impact; moderate consequence | Pricing recommendations, fraud detection, hiring screening |
| 3 | Rights-critical, safety-critical, or irreversible harm | Medical diagnosis assist, loan decisions, legal advice, safety systems |
Dimension D3: Autonomy
Question: How hands-off is the system?
| Score | Description | Examples |
| 0 | Reference/advisory only; no actions taken | Research assistant, comparison tool, data explorer |
| 1 | Draft generation; human approval required before action | Email drafter, report generator, code suggestion |
| 2 | Automated actions with notification; human can intervene | Auto-categorization with review queue, smart routing |
| 3 | Fully autonomous; minimal human oversight | Autonomous trading, auto-moderation, dynamic pricing |
Dimension D4: Data Sensitivity
Question: How sensitive are inputs/outputs?
| Score | Description | Examples |
| 0 | Public information only | News summarizer, public data analyzer, general Q&A |
| 1 | Internal business data; non-sensitive | Sales data, product catalogs, project documentation |
| 2 | PII, confidential, or proprietary data | Customer data, employee records, trade secrets |
| 3 | Regulated data (health, financial, biometric) | PHI/HIPAA, payment card data, biometric identifiers |
Dimension D5: Model/Change Risk
Question: How capable and dynamic is the system?
| Score | Description | Examples |
| 0 | Stable, narrow-task model; infrequent updates | Simple classifier, static recommendation engine |
| 1 | Established model; monthly updates | Standard NLP, computer vision with periodic retraining |
| 2 | Advanced capabilities; weekly updates or tool use | LLM with retrieval, multi-step reasoning, API calls |
| 3 | Frontier model; rapid iteration, multi-modal, agentic | Latest GPT/Claude with tools, autonomous agents, multi-modal |
Dimension D6: Misuse Potential
Question: What's the abuse value to attackers?
| Score | Description | Examples |
| 0 | Low value to attackers; limited abuse scenarios | Internal productivity tool, low-stakes recommendations |
| 1 | Moderate value; requires effort to abuse | Content generation, simple chatbot |
| 2 | High value; clear abuse paths exist | Code generation, advanced research assistant |
| 3 | Critical abuse potential; scalable harm scenarios | Disinformation generation, exploit creation, impersonation at scale |
8.2 Calculating the Risk Tier
Step 1: Score Each Dimension
Assign 0-3 for each of the 6 dimensions (D1-D6)
Document brief justification for each score
Be honest and conservative (when in doubt, score higher)
Step 2: Calculate Total Score
Sum all dimension scores: Total = D1 + D2 + D3 + D4 + D5 + D6
Range: 0 (lowest risk) to 18 (highest risk)
Step 3: Determine Initial Tier
| Total Score | Initial Tier |
| 0 - 6 | Tier 1 (Low) |
| 7 - 12 | Tier 2 (Medium) |
| 13 - 18 | Tier 3 (High) |
Step 4: Apply Automatic Tier 3 Overrides
Even if the calculated score is lower, the system is automatically Tier 3 if ANY of the following are true:
☐ Public-facing AND touches personal/regulated data
☐ Provides rights/safety-critical advice or materially influences determinations (employment, credit, legal, medical, safety)
☐ Autonomous actions can materially impact users or systems without human review
Step 5: Final Tier Assignment
Final Tier = Higher of (Calculated Tier, Override Tier)
Document which override triggered (if applicable)
8.3 Tiering Examples
Example 1: Internal Meeting Summarizer
D1 (Exposure): 0 - Internal team only
D2 (Impact): 0 - Convenience only
D3 (Autonomy): 1 - Generates draft; user reviews
D4 (Data): 1 - Internal non-sensitive meetings
D5 (Model/Change): 1 - Stable model, monthly updates
D6 (Misuse): 0 - Low abuse value
Total: 3 → Tier 1
Override Check: None apply
Final: Tier 1
Example 2: Customer Support Chatbot
D1 (Exposure): 3 - Public-facing
D2 (Impact): 1 - Convenience/productivity
D3 (Autonomy): 1 - Drafts responses; agent reviews
D4 (Data): 2 - Customer PII
D5 (Model/Change): 2 - LLM with retrieval, weekly updates
D6 (Misuse): 1 - Moderate abuse risk
Total: 10 → Tier 2
Override Check: Public + PII → Tier 3 Override
Final: Tier 3
Example 3: Internal HR Policy Assistant
D1 (Exposure): 1 - Organization-wide internal
D2 (Impact): 2 - Guidance affects decisions
D3 (Autonomy): 0 - Reference only
D4 (Data): 2 - Employee data, policies
D5 (Model/Change): 2 - LLM with RAG
D6 (Misuse): 1 - Moderate (wrong guidance risk)
Total: 8 → Tier 2
Override Check: None apply
Final: Tier 2
8.4 Re-Tiering Requirements
Re-tier the AI system when:
Material change to intended use (new user population, new use cases)
Significant architecture change (new model, added tools/capabilities)
Data sensitivity change (now processing regulated data)
Autonomy increase (removing human review steps)
Deployment change (internal → public)
Incident revealing higher risk than initially assessed
Annual review (at minimum, confirm tier remains accurate)
9. Required Artifacts and Approvals by Tier
9.1 Artifact Summary Table
| Tier | Approvals Required | Required Artifacts | Stop-Ship Gate |
| Tier 1 (Low) | Business Owner + Engineering Owner | • System Card<br>• Data Handling Note<br>• Light TEVV<br>• Logging/Retention plan | No launch without artifacts and sign-offs |
| Tier 2 (Medium) | ARB delegate + Security + Privacy/Legal + Domain Lead | Tier 1 artifacts PLUS:<br>• Risk Analysis<br>• Threat Model<br>• TEVV Report (full)<br>• Oversight specification<br>• Monitoring thresholds | No launch without passing TEVV thresholds and signed reviews |
| Tier 3 (High) | Full ARB + Security + Privacy/Legal + Executive Risk Owner | Tier 2 artifacts PLUS:<br>• Independent validation<br>• Provenance Plan (GenAI)<br>• Red-team testing + retest<br>• Disclosure readiness<br>• Rollback/kill switch<br>• Change control | No launch without Tier 3 evidence; residual risk requires executive acceptance |
9.2 Tier 1 Requirements (Low Risk)
Objective: Ensure basic safety, transparency, and monitoring for low-risk internal tools.
Required Artifacts:
System Card (Template: Section 11.2)
Purpose and intended use
Model details
Data handling (inputs, outputs, logging)
Human oversight approach
Known limitations
Monitoring plan
Data Handling Note
What data is collected
Retention period
Access controls
What is NOT logged (PII redaction if applicable)
Light TEVV
20-100 test cases (depends on complexity)
Basic accuracy/quality threshold
Simple failure mode identification
Pass/fail determination
Logging and Retention Plan
What gets logged (metadata, errors, usage)
Retention period (typically 30-90 days)
Access controls
Approval Process:
Business Owner signs off on use case and limitations
Engineering Owner signs off on technical implementation
AI Risk Lead validates completeness
Timeline: 3-5 business days typical
Post-Launch:
Review 30-60 days after launch
Monitor for issues
Update tier if risk profile changes
9.3 Tier 2 Requirements (Medium Risk)
Objective: Comprehensive risk assessment, testing, and ongoing monitoring.
Required Artifacts (in addition to Tier 1):
Risk Analysis (Template: Section 11.6)
Context summary
Risk register with severity, likelihood, controls
GenAI-specific risks (if applicable)
Risk treatment decisions
Linkage to TEVV and monitoring
Threat Model (Template: Section 11.4)
System diagram and trust boundaries
Assets to protect
Threat actors and scenarios
Controls checklist
Open risks and mitigations
Full TEVV Report (Template: Section 11.3)
Functional/quality evaluation with metrics
Safety evaluation
Reliability/robustness testing
Fairness/bias testing (if applicable)
Privacy leakage tests
Security tests (prompt injection, etc.)
Results with pass/fail against thresholds
Retest results after mitigations
Oversight Specification
Human review points
Escalation paths
Override mechanisms
Quality sampling plan
Monitoring Thresholds
Performance metrics and SLOs
Alert thresholds
On-call/escalation
Dashboards
Approval Process:
All Tier 1 approvers PLUS:
Security reviews threat model and security testing
Privacy/Legal reviews data handling and compliance
Domain Lead reviews quality standards and TEVV
ARB (or delegate) reviews complete package
Timeline: 1-2 weeks typical
Post-Launch:
Review 30 days after launch
Quarterly reviews thereafter
Incident-driven reviews as needed
9.4 Tier 3 Requirements (High Risk)
Objective: Maximum assurance through independent validation, red-team testing, and executive oversight.
Required Artifacts (in addition to Tier 2):
Independent Validation
Third-party or independent internal team
Validates TEVV results
Validates risk assessment
Validates control implementation
Issues validation report
Provenance & Transparency Plan (GenAI - Template: Section 11.5)
Content requiring provenance tracking
User-facing transparency (labels, disclosures)
Technical provenance mechanisms
Audit trail components
Operational process
Red-Team Testing + Retest
Comprehensive adversarial scenarios
Jailbreak attempts
Prompt injection testing
Data exfiltration attempts
Social engineering scenarios
Initial test results
Mitigations implemented
Retest showing improvements
Disclosure Readiness
Triggers for disclosure (incident types, severity)
Internal notification process
External notification process
Regulatory notification requirements
Communication templates
Responsible parties
Rollback/Kill Switch
Technical mechanism to disable system
Rollback to previous version capability
Decision authority for activation
Testing verification
Communication plan
Change Control Process
What changes require re-approval
Re-testing requirements
Approval workflow for changes
Rollback plan for failed changes
Approval Process:
All Tier 2 approvers PLUS:
Full ARB review (not delegate)
Executive Risk Owner reviews and accepts residual risk
Timeline: 2-4 weeks typical
Post-Launch:
Review 2 weeks after launch
Monthly reviews for first 3 months
Quarterly reviews thereafter
Executive dashboard reporting
10. Approval Workflow
10.1 Workflow Overview
The approval workflow maps to NIST AI RMF functions:
Submission (GOVERN) - Intake Form with initial artifacts
Triage & Tiering (GOVERN) - AI Risk Lead scores tier
Context & Data Review (MAP) - System Card approval
Evaluation Review (MEASURE) - TEVV validation
Operational Readiness (MANAGE) - Monitoring and rollback
Approval Decision (GOVERN) - Sign-offs obtained
Launch & Post-Launch (MANAGE) - Deploy with monitoring
Change Control (GOVERN/MEASURE) - Material changes
10.2 Stage 1: Submission (GOVERN)
Owner: Business Owner + Engineering Owner
Activities:
Complete AI Intake Form (Template: Section 11.1)
Attach initial System Card
Provide initial risk tier self-assessment
Submit to AI Risk Lead
Quality Gates:
All intake form fields completed
Use case clearly described
Initial artifacts attached
Owners identified
Timeline: Submit when ready to begin governance review
Outcome: Intake accepted or returned for completion
10.3 Stage 2: Triage & Tiering (GOVERN)
Owner: AI Risk Lead
Activities:
Review intake submission
Validate or adjust risk tier using rubric (Section 8)
Check for automatic Tier 3 overrides
Assign required artifacts and reviewers
Create tracking ticket
Notify team of tier and requirements
Quality Gates:
Risk tier justified with dimension scores
Required artifacts list generated
Reviewers assigned
Timeline communicated
Timeline: 1-2 business days
Outcome: System tiered; requirements documented; review initiated
10.4 Stage 3: Context & Data Review (MAP)
Owner: Business Owner + Engineering Owner
Reviewers: Privacy/Legal (Tier 2+)
Activities:
Finalize System Card
Document intended use, users, constraints
Map data flows (inputs, outputs, storage)
Identify prohibited uses
Document human oversight approach
Privacy/Legal reviews data handling
Quality Gates:
System Card complete and approved
Data handling compliant with privacy policies
Prohibited uses explicitly documented
Human oversight defined
Timeline: 3-5 days (Tier 1); 1 week (Tier 2-3)
Outcome: System Card approved; data handling validated
10.5 Stage 4: Evaluation Review (MEASURE)
Owner: Engineering Owner + Domain Lead
Reviewers: Security (Tier 2+), Domain Lead (Tier 2+)
Activities:
Execute TEVV testing plan
Security testing (prompt injection, etc.)
Fairness/bias testing (if applicable)
Domain correctness validation
Document results and mitigations
Retest after fixes
Independent validation (Tier 3)
Red-team testing (Tier 3)
Quality Gates:
All TEVV tests executed
Results meet or exceed thresholds
Security testing passed
Domain quality validated
Mitigations implemented and retested
Independent validation report (Tier 3)
Red-team retest passed (Tier 3)
Timeline: 1 week (Tier 1); 2 weeks (Tier 2); 3-4 weeks (Tier 3)
Outcome: TEVV approved; system validated ready for deployment
10.6 Stage 5: Operational Readiness (MANAGE)
Owner: Engineering Owner
Activities:
Configure production monitoring
Set alert thresholds
Test rollback mechanism
Document incident response
Create operational runbook
Test kill switch (Tier 3)
Finalize Provenance Plan (GenAI Tier 3)
Quality Gates:
Monitoring operational before launch
Alerts configured and tested
Rollback tested successfully
Incident response documented
On-call identified
Kill switch verified (Tier 3)
Timeline: 3-5 days
Outcome: System operationally ready; monitoring live
10.7 Stage 6: Approval Decision (GOVERN)
Owner: ARB (Tier 2-3) or Business/Engineering Owners (Tier 1)
Activities:
Review complete artifact package
Validate all quality gates passed
Assess residual risks
Make launch/no-launch decision
Executive risk acceptance (Tier 3)
Document approvals
Quality Gates:
All required artifacts complete
All reviewers signed off
TEVV thresholds met
Monitoring ready
Residual risks accepted (Tier 3)
Decision Options:
Approve: Launch authorized
Approve with Conditions: Launch with specific constraints
Defer: Additional work required; resubmit
Reject: Do not proceed
Timeline: 1-2 days (Tier 1); 3-5 days (Tier 2); 1 week (Tier 3)
Outcome: Launch authorization or remediation requirements
10.8 Stage 7: Launch & Post-Launch Review (MANAGE)
Owner: Engineering Owner
Activities:
Deploy to production
Verify monitoring operational
Communicate to users
Monitor closely during initial period
Conduct post-launch review
Post-Launch Review Cadence:
Tier 1: 30-60 days after launch
Tier 2: 30 days after launch
Tier 3: 2 weeks, then monthly for 3 months, quarterly thereafter
Post-Launch Review Agenda:
Review monitoring metrics vs. TEVV results
Assess incidents or issues
Validate risk tier remains accurate
Identify improvements
Update artifacts as needed
Timeline: Ongoing
Outcome: System in production with active monitoring and governance
10.9 Stage 8: Change Control (GOVERN/MEASURE)
Trigger: Material changes to the AI system
Material Changes Include:
New use cases or user populations
Model updates (major version, fine-tuning changes)
New data sources or types
Architecture changes (adding tools, changing RAG)
Autonomy changes (removing human review)
Increased deployment scope
Change Control Process:
Assess if change is material
If material: Re-tier using rubric
Determine required re-testing (partial or full TEVV)
Execute required testing
Update artifacts
Obtain re-approval per tier
Deploy with monitoring
Non-Material Changes:
Bug fixes with no behavior change
Infrastructure updates with no model change
Monitoring improvements
Documentation updates
Timeline: Depends on extent of change and testing required
Outcome: Changed system re-approved or change deferred
10.10 Workflow SLA Summary
| Stage | Tier 1 | Tier 2 | Tier 3 |
| Submission | 1 day | 1 day | 1 day |
| Triage & Tiering | 1-2 days | 1-2 days | 1-2 days |
| Context Review | 3-5 days | 5-7 days | 5-7 days |
| Evaluation Review | 5-7 days | 10-14 days | 15-21 days |
| Operational Readiness | 3-5 days | 3-5 days | 5-7 days |
| Approval Decision | 1-2 days | 3-5 days | 5-7 days |
| TOTAL END-TO-END | 2-3 weeks | 4-6 weeks | 6-8 weeks |
Note: Timelines assume artifacts are complete and testing passes on first attempt. Rework extends timeline.
11. Forms and Templates
11.1 AI Intake Form (Template)
Copy/paste the following form for every AI system submission.
Copy
================================================================================
AI SYSTEM INTAKE FORM
================================================================================
A) REQUEST METADATA
-------------------
System name: _________________________________
Business Owner: _________________________________
Engineering Owner: _________________________________
Domain / Product: _________________________________
GenAI? (Y/N): _________________________________
Vendor / Internal / Hybrid: _________________________________
User population (team / org / partner / public): _________________________________
Regions impacted: _________________________________
Target launch date: _________________________________
B) SYSTEM DESCRIPTION (MAP)
---------------------------
What does it do (plain language)?
_________________________________________________________________
_________________________________________________________________
Intended uses:
_________________________________________________________________
_________________________________________________________________
Explicit prohibited uses:
_________________________________________________________________
_________________________________________________________________
Where does it appear (workflow/UI/API)?
_________________________________________________________________
Human oversight (who reviews, when)?
_________________________________________________________________
Does it take actions or only generate content?
_________________________________________________________________
C) MODEL DETAILS (MAP)
----------------------
Model(s) + version(s):
_________________________________________________________________
Prompting approach / tools / retrieval sources:
_________________________________________________________________
_________________________________________________________________
Fine-tuning? (Y/N; describe):
_________________________________________________________________
Update frequency:
_________________________________________________________________
Fallback behavior:
_________________________________________________________________
D) DATA & PRIVACY (MAP)
-----------------------
Inputs include (PII/regulated/confidential?):
_________________________________________________________________
_________________________________________________________________
Outputs may contain (PII/sensitive?):
_________________________________________________________________
Logging (what is stored; what is excluded):
_________________________________________________________________
_________________________________________________________________
Retention period:
_________________________________________________________________
Access controls (who can access prompts/outputs/logs?):
_________________________________________________________________
Vendor data usage terms reviewed? (Y/N + link):
_________________________________________________________________
E) RISK SCORING (GOVERN)
------------------------
D1 Exposure (0-3) + justification:
Score: _____ | Justification: _________________________________
D2 Impact (0-3) + justification:
Score: _____ | Justification: _________________________________
D3 Autonomy (0-3) + justification:
Score: _____ | Justification: _________________________________
D4 Data sensitivity (0-3) + justification:
Score: _____ | Justification: _________________________________
D5 Model/change risk (0-3) + justification:
Score: _____ | Justification: _________________________________
D6 Misuse potential (0-3) + justification:
Score: _____ | Justification: _________________________________
Total score (0-18): _____
Calculated Tier: _____
Tier 3 Override Check:
☐ Public-facing AND touches personal/regulated data
☐ Provides rights/safety-critical advice or materially influences determinations
☐ Autonomous actions can materially impact users or systems
Final Tier Assignment: _____
F) EVIDENCE CHECKLIST (MEASURE/MANAGE) - LINKS
-----------------------------------------------
System Card: _________________________________
TEVV Report: _________________________________
Threat Model (Tier 2+): _________________________________
Risk Analysis (Tier 2+): _________________________________
Provenance Plan (GenAI Tier 3+): _________________________________
Independent validation (Tier 3): _________________________________
Monitoring dashboard link: _________________________________
Rollback plan link: _________________________________
G) SIGN-OFFS
------------
Business Owner: _______________________ Date: __________
Engineering Owner: _______________________ Date: __________
Security: _______________________ Date: __________
Privacy/Legal: _______________________ Date: __________
Domain Lead: _______________________ Date: __________
ARB Chair/Delegate: _______________________ Date: __________
Executive Risk Owner (Tier 3): _______________________ Date: __________
================================================================================
11.2 System Card (Template)
Use this template for every AI system. Keep it 1-3 pages for most systems; add appendices if needed.
Copy
================================================================================
AI SYSTEM CARD
================================================================================
DOCUMENT METADATA
-----------------
Owner: _________________________________
Reviewers: _________________________________
Version: _________________________________
Date: _________________________________
Tier: _________________________________
1) SYSTEM OVERVIEW
------------------
System name / ID: _________________________________
Type (GenAI/ML/etc.): _________________________________
Purpose (1-2 sentences):
_________________________________________________________________
_________________________________________________________________
Users: _________________________________
Where used: _________________________________
Intended outputs: _________________________________
Explicit non-intended uses:
_________________________________________________________________
2) MODEL AND ARCHITECTURE
-------------------------
Model(s) and versions: _________________________________
Provider (internal/vendor): _________________________________
Retrieval/tools/plugins used: _________________________________
Prompting strategy:
_________________________________________________________________
Fine-tuning (data/method/frequency):
_________________________________________________________________
Dependencies: _________________________________
3) DATA HANDLING
----------------
Inputs (data types; PII/regulated?):
_________________________________________________________________
Outputs (sensitive content risk?):
_________________________________________________________________
Logging (what is logged; what is never logged):
_________________________________________________________________
Retention: _________________________________
Access controls: _________________________________
Data minimization controls:
_________________________________________________________________
Vendor data usage constraints:
_________________________________________________________________
4) HUMAN OVERSIGHT AND OPERATING CONSTRAINTS
--------------------------------------------
Human-in-the-loop review points:
_________________________________________________________________
User warnings/instructions:
_________________________________________________________________
Safety constraints / refusal behavior:
_________________________________________________________________
Autonomy boundaries (allowed actions):
_________________________________________________________________
5) RISKS AND LIMITATIONS
------------------------
Known failure modes:
_________________________________________________________________
_________________________________________________________________
Known limitations:
_________________________________________________________________
_________________________________________________________________
Out-of-scope decisions:
_________________________________________________________________
6) MONITORING AND ROLLBACK
--------------------------
Monitoring metrics:
_________________________________________________________________
Alert thresholds:
_________________________________________________________________
On-call/owner: _________________________________
Rollback/kill switch steps:
_________________________________________________________________
_________________________________________________________________
Post-launch review schedule: _________________________________
7) CHANGE CONTROL
-----------------
Changes requiring re-approval:
_________________________________________________________________
Re-test requirements:
_________________________________________________________________
================================================================================
11.3 TEVV Report (Template)
Use this template for Tier 2-3. Tier 1 may use a reduced ("light TEVV") version.
Copy
================================================================================
TEVV REPORT (Testing, Evaluation, Validation, Verification)
================================================================================
DOCUMENT METADATA
-----------------
Owner: _________________________________
Version/Date: _________________________________
System Name: _________________________________
Tier: _________________________________
1) EVALUATION OBJECTIVES
------------------------
What you must prove before launch:
_________________________________________________________________
_________________________________________________________________
_________________________________________________________________
2) TEST PLAN SUMMARY
--------------------
What was tested and how:
A. FUNCTIONAL/QUALITY EVALUATION
Task suite: _________________________________
Rubric/metrics: _________________________________
Thresholds: _________________________________
B. SAFETY EVALUATION
Disallowed content categories: _________________________________
Refusal requirements: _________________________________
Testing approach: _________________________________
C. RELIABILITY/ROBUSTNESS
Regression tests: _________________________________
Edge cases: _________________________________
Load/stress testing: _________________________________
D. FAIRNESS/BIAS (if applicable)
Protected attributes: _________________________________
Subgroups tested: _________________________________
Fairness metrics: _________________________________
Thresholds: _________________________________
E. PRIVACY LEAKAGE TESTS
PII echo tests: _________________________________
RAG leakage tests: _________________________________
Memorization checks: _________________________________
F. SECURITY TESTS
Prompt injection: _________________________________
Tool abuse: _________________________________
Access bypass: _________________________________
Data exfiltration: _________________________________
G. ADVERSARIAL/RED-TEAM TESTING (Tier 3 required)
Attack scenarios: _________________________________
Testing team: _________________________________
Duration: _________________________________
3) RESULTS
----------
For each section above, document:
A. FUNCTIONAL/QUALITY RESULTS
Cases tested: _________________________________
Pass rate: _________________________________
Top failure modes:
_________________________________________________________________
Severity: _________________________________
Mitigations:
_________________________________________________________________
Retest results: _________________________________
B. SAFETY RESULTS
Cases tested: _________________________________
Refusal rate: _________________________________
False refusals: _________________________________
Harmful outputs: _________________________________
Mitigations:
_________________________________________________________________
Retest results: _________________________________
C. RELIABILITY/ROBUSTNESS RESULTS
Regression pass rate: _________________________________
Edge case performance: _________________________________
Load test results: _________________________________
Mitigations:
_________________________________________________________________
D. FAIRNESS/BIAS RESULTS
Subgroup performance:
- Group A: _________________________________
- Group B: _________________________________
- Group C: _________________________________
Disparity metrics: _________________________________
Mitigations:
_________________________________________________________________
Retest results: _________________________________
E. PRIVACY LEAKAGE RESULTS
PII echo rate: _________________________________
RAG leakage incidents: _________________________________
Memorization detected: _________________________________
Mitigations:
_________________________________________________________________
F. SECURITY RESULTS
Prompt injection resistance: _________________________________
Tool abuse attempts blocked: _________________________________
Access control bypass: _________________________________
Mitigations:
_________________________________________________________________
Retest results: _________________________________
G. ADVERSARIAL/RED-TEAM RESULTS (Tier 3)
Total attack scenarios: _________________________________
Successful attacks (before mitigation): _________________________________
Attack categories:
_________________________________________________________________
Critical findings:
_________________________________________________________________
Mitigations implemented:
_________________________________________________________________
Retest results: _________________________________
Residual vulnerabilities:
_________________________________________________________________
4) LAUNCH DECISION
------------------
Meets thresholds? (Y/N): _________________________________
Residual risks:
_________________________________________________________________
_________________________________________________________________
Risk acceptance required? (Y/N): _________________________________
If yes, approved by: _________________________________
Launch recommendation: ☐ Approve ☐ Approve with conditions ☐ Defer ☐ Reject
Conditions (if applicable):
_________________________________________________________________
5) MONITORING LINKAGE
---------------------
Which metrics detect which risks post-launch:
Risk: _________________ → Metric: _________________ → Threshold: _______
Risk: _________________ → Metric: _________________ → Threshold: _______
Risk: _________________ → Metric: _________________ → Threshold: _______
================================================================================
11.4 Threat Model (Template)
Copy
================================================================================
AI THREAT MODEL
================================================================================
DOCUMENT METADATA
-----------------
Owner: _________________________________
Version/Date: _________________________________
System Name: _________________________________
Tier: _________________________________
1) SYSTEM DIAGRAM
-----------------
(Text description or attach diagram)
Components:
_________________________________________________________________
_________________________________________________________________
Trust boundaries:
_________________________________________________________________
External interfaces:
_________________________________________________________________
2) ASSETS TO PROTECT
--------------------
☐ Sensitive data (describe): _________________________________
☐ Retrieval corpora/knowledge bases: _________________________________
☐ Credentials/API keys: _________________________________
☐ Tools/actions/capabilities: _________________________________
☐ System prompts/configuration: _________________________________
☐ Logs/audit trails: _________________________________
☐ Model weights/parameters: _________________________________
☐ Other: _________________________________
3) THREAT ACTORS AND MOTIVATIONS
--------------------------------
☐ External attacker (motivation: data theft, disruption, financial gain)
☐ Insider threat (motivation: curiosity, malice, financial gain)
☐ Curious user (motivation: exploration, testing limits)
☐ Competitor (motivation: intelligence gathering, sabotage)
☐ Automated bots (motivation: scraping, abuse at scale)
☐ Other: _________________________________
4) THREAT SCENARIOS
-------------------
For each threat, document: Attack path | Impact | Likelihood | Controls | Residual risk
REQUIRED SCENARIOS:
A. PROMPT INJECTION
Attack path: _________________________________
Impact: _________________________________
Likelihood: _________________________________
Controls: _________________________________
Residual risk: _________________________________
B. DATA EXFILTRATION
Attack path: _________________________________
Impact: _________________________________
Likelihood: _________________________________
Controls: _________________________________
Residual risk: _________________________________
C. UNAUTHORIZED TOOL INVOCATION
Attack path: _________________________________
Impact: _________________________________
Likelihood: _________________________________
Controls: _________________________________
Residual risk: _________________________________
D. INDIRECT PROMPT INJECTION (if retrieval/RAG)
Attack path: _________________________________
Impact: _________________________________
Likelihood: _________________________________
Controls: _________________________________
Residual risk: _________________________________
E. DENIAL OF SERVICE / COST BLOW-UP
Attack path: _________________________________
Impact: _________________________________
Likelihood: _________________________________
Controls: _________________________________
Residual risk: _________________________________
F. IMPERSONATION / DISINFORMATION
Attack path: _________________________________
Impact: _________________________________
Likelihood: _________________________________
Controls: _________________________________
Residual risk: _________________________________
G. JAILBREAKS / SAFETY BYPASS
Attack path: _________________________________
Impact: _________________________________
Likelihood: _________________________________
Controls: _________________________________
Residual risk: _________________________________
H. ACCESS CONTROL BYPASS
Attack path: _________________________________
Impact: _________________________________
Likelihood: _________________________________
Controls: _________________________________
Residual risk: _________________________________
5) CONTROLS CHECKLIST
---------------------
☐ Input allowlisting/validation
☐ Authorization checks (RBAC/ABAC)
☐ Rate limiting (per user, per IP, global)
☐ Content filtering (input and output)
☐ Logging and auditing
☐ Secrets management (no secrets in prompts)
☐ RAG access controls (least privilege)
☐ Sandboxing/isolation for tools
☐ Output validation before action
☐ Anomaly detection
☐ Session management
☐ Encryption (in transit, at rest)
Describe implementation:
_________________________________________________________________
_________________________________________________________________
6) OPEN RISKS AND REQUIRED MITIGATIONS
---------------------------------------
MUST-FIX BEFORE LAUNCH:
_________________________________________________________________
_________________________________________________________________
ACCEPT (with justification):
_________________________________________________________________
_________________________________________________________________
================================================================================
11.5 Provenance & Transparency Plan (Template)
Required for Tier 3 GenAI; recommended for other GenAI that produces distributable content.
Copy
================================================================================
PROVENANCE & TRANSPARENCY PLAN
================================================================================
DOCUMENT METADATA
-----------------
Owner: _________________________________
Version/Date: _________________________________
System Name: _________________________________
1) CONTENT REQUIRING PROVENANCE
-------------------------------
Content types needing provenance tracking:
☐ Text
☐ Images
☐ Audio
☐ Video
☐ Code
☐ Other: _________________________________
Where content appears:
_________________________________________________________________
Distribution scope (internal / partner / public):
_________________________________________________________________
2) USER-FACING TRANSPARENCY REQUIREMENTS
----------------------------------------
Labels/disclosures:
_________________________________________________________________
Tooltip/help text:
_________________________________________________________________
Limitations communicated to users:
_________________________________________________________________
Where disclosures appear:
_________________________________________________________________
3) PROVENANCE DATA CAPTURED (AUDIT TRAIL)
------------------------------------------
☐ Model name and version
☐ Prompt template version
☐ Retrieval source IDs (if RAG)
☐ Timestamp (creation, modification)
☐ Tool calls made (if agentic)
☐ Post-edit indicator (human modifications)
☐ Approval status (if workflow)
☐ User ID (if tracked)
☐ Other: _________________________________
4) TECHNICAL MECHANISMS
-----------------------
Metadata embedding:
_________________________________________________________________
Internal content IDs:
_________________________________________________________________
Watermark/signature (if used):
_________________________________________________________________
Downstream survivability limitations:
_________________________________________________________________
Storage location for provenance data:
_________________________________________________________________
5) OPERATIONAL PROCESS
----------------------
Access to provenance logs:
_________________________________________________________________
Audit cadence:
_________________________________________________________________
Dispute handling (if user questions AI content):
_________________________________________________________________
Retention constraints:
_________________________________________________________________
6) ACCEPTANCE CRITERIA
----------------------
Labeling coverage target: _______%
Provenance completeness target: _______%
Failure handling (what happens if provenance unavailable):
_________________________________________________________________
Launch readiness checklist:
☐ Labeling tested and verified
☐ Provenance capture tested
☐ Audit trail accessible
☐ User disclosure implemented
☐ Failure modes handled gracefully
================================================================================
11.6 Risk Analysis (Template)
Use this template for Tier 2-3. Focus is risk identification and treatment.
Copy
================================================================================
AI RISK ANALYSIS
================================================================================
DOCUMENT METADATA
-----------------
Owner: _________________________________
Version/Date: _________________________________
System Name: _________________________________
Tier: _________________________________
1) CONTEXT SUMMARY
------------------
Intended use:
_________________________________________________________________
Users:
_________________________________________________________________
Deployment environment:
_________________________________________________________________
Data types processed:
_________________________________________________________________
Autonomy level:
_________________________________________________________________
What "bad outcome" looks like:
_________________________________________________________________
_________________________________________________________________
2) RISK REGISTER
----------------
For each risk, document:
Risk ID | Risk statement | Type | Harmed parties | Severity | Likelihood | Inherent rating | Existing controls | Mitigations | Owner | Due date | Residual rating | Acceptance criteria
EXAMPLE FORMAT:
Risk-001
Statement: Hallucinated policy guidance leads to incorrect employee decisions
Type: Misinformation
Harmed parties: Employees, organization
Severity: High
Likelihood: Medium
Inherent rating: 12 (High x Medium)
Existing controls: Citation requirement, source linking
Mitigations: Post-generation fact-check, hallucination detection, user warnings
Owner: Engineering Owner
Due date: Before launch
Residual rating: 6 (Medium x Low)
Acceptance: Hallucination rate < 3%, citation coverage > 95%
[Add 5-15 risks depending on system complexity]
3) GENAI-SPECIFIC RISK CATEGORIES (if GenAI)
--------------------------------------------
Assess each category:
A. HALLUCINATION/MISINFORMATION
Risk present? ☐ Yes ☐ No
Description: _________________________________
Severity: _________________________________
Mitigations: _________________________________
B. PROMPT INJECTION
Risk present? ☐ Yes ☐ No
Description: _________________________________
Severity: _________________________________
Mitigations: _________________________________
C. DATA LEAKAGE
Risk present? ☐ Yes ☐ No
Description: _________________________________
Severity: _________________________________
Mitigations: _________________________________
D. DISALLOWED CONTENT GENERATION
Risk present? ☐ Yes ☐ No
Description: _________________________________
Severity: _________________________________
Mitigations: _________________________________
E. IMPERSONATION/MANIPULATION
Risk present? ☐ Yes ☐ No
Description: _________________________________
Severity: _________________________________
Mitigations: _________________________________
F. PROVENANCE GAPS
Risk present? ☐ Yes ☐ No
Description: _________________________________
Severity: _________________________________
Mitigations: _________________________________
G. MODEL UPDATE REGRESSIONS
Risk present? ☐ Yes ☐ No
Description: _________________________________
Severity: _________________________________
Mitigations: _________________________________
4) RISK TREATMENT DECISION (FOR HIGH/CRITICAL RISKS)
----------------------------------------------------
For each high/critical risk:
Risk ID: _________
Treatment: ☐ Mitigate ☐ Avoid ☐ Transfer ☐ Accept
Justification: _________________________________
Approvers: _________________________________
5) LINKAGE TO TEVV AND MONITORING
---------------------------------
Test coverage for each high risk:
Risk: _________ → TEVV test: _________ → Pass threshold: _________
Monitoring metrics for drift detection:
Risk: _________ → Metric: _________ → Alert threshold: _________
6) DISCLOSURE READINESS (TIER 3)
--------------------------------
Triggers for disclosure:
_________________________________________________________________
Internal notification owners:
_________________________________________________________________
External notification criteria:
_________________________________________________________________
Regulatory notification requirements:
_________________________________________________________________
Communication templates prepared? ☐ Yes ☐ No
================================================================================
12. Global Regulatory Landscape 2024-2025
12.1 Overview
The AI regulatory landscape has evolved significantly in 2024-2025, with major jurisdictions implementing binding requirements:
Key Developments:
EU AI Act entered into force August 2024, with phased implementation through 2027
US Executive Order 14110 (October 2023) continues to drive federal AI governance
UK AI regulation following principles-based approach with sector-specific rules
China enforcing Generative AI regulations since August 2023
State-level US laws (California, Colorado, New York) creating compliance patchwork
12.2 EU AI Act (Regulation 2024/1689)
Status: In force August 2, 2024; phased implementation
Scope: AI systems placed on EU market or affecting EU persons
Key Classifications:
Prohibited AI (banned):
Subliminal manipulation causing harm
Exploitation of vulnerabilities (age, disability)
Social scoring by governments
Real-time biometric identification in public (with narrow exceptions)
Emotion recognition in workplace/education (with exceptions)
High-Risk AI (heavy requirements):
Employment/HR systems
Credit scoring and creditworthiness
Education/exam scoring
Law enforcement risk assessment
Border control biometric verification
Critical infrastructure safety components
General Purpose AI Models (GPAI):
Transparency and documentation requirements
Copyright compliance policy
Systemic risk GPAI (>10^25 FLOPs): Enhanced requirements including adversarial testing, incident reporting
Compliance Timeline:
February 2, 2025: Prohibited practices ban
August 2, 2025: GPAI obligations
August 2, 2026: High-risk requirements for new systems
August 2, 2027: High-risk requirements for existing systems
Penalties:
Up to €35M or 7% global revenue (prohibited AI)
Up to €15M or 3% global revenue (obligations breach)
How This Framework Helps:
Risk tiering aligns with EU risk levels
Tier 3 requirements meet high-risk AI obligations
TEVV process satisfies conformity assessment needs
Documentation templates support EU transparency requirements
12.3 GDPR (EU) and UK GDPR
Impact on AI:
Data minimization: Collect only necessary data for AI training/inference
Purpose limitation: Use data only for specified AI purposes
Lawful basis: Require consent, contract, or legitimate interest for AI processing
Article 22: Right not to be subject to solely automated decisions with legal/significant effects
Data Protection Impact Assessment (DPIA): Required for high-risk AI processing personal data
Transparency: Inform data subjects about AI use
How This Framework Helps:
Privacy/Legal review ensures GDPR compliance
System Card documents data handling and lawful basis
TEVV includes privacy leakage testing
Risk Analysis identifies GDPR risks
12.4 CCPA/CPRA (California)
Requirements for AI:
Notice: Inform consumers about automated decision-making
Access: Right to know if AI used in decisions affecting them
Opt-out: Right to opt out of sale/sharing of personal information for AI training
Correction: Right to correct inaccurate personal information
Automated decision-making technology: Additional disclosure requirements
How This Framework Helps:
Privacy review covers CCPA requirements
Transparency controls enable consumer notice
Monitoring supports access/correction requests
12.5 US Federal (Executive Order 14110)
Key Requirements:
Safety testing: For models trained with >10^26 FLOPs
Red-team testing: Before public release
Information sharing: Security vulnerabilities with government
Equity and civil rights: Prevent algorithmic discrimination
Privacy: Privacy-preserving techniques
Agency-specific guidance: OMB memo for federal agencies
How This Framework Helps:
Red-team testing required for Tier 3
Fairness/bias testing addresses equity requirements
Threat modeling covers security requirements
Provenance supports transparency
12.6 Sector-Specific Regulations
Healthcare (HIPAA, FDA):
HIPAA: PHI protection in AI systems
FDA: AI/ML Software as Medical Device (SaMD) regulations
Requirements: Clinical validation, transparency, ongoing monitoring
Financial Services (SOX, FCRA, ECOA):
Fair lending: Equal Credit Opportunity Act applies to AI credit decisions
Model risk management: Federal Reserve SR 11-7 for model validation
Explainability: Adverse action notices require reason codes
Employment (EEOC, Labor Laws):
Anti-discrimination: Title VII, ADA apply to AI hiring/HR tools
EEOC guidance: Algorithmic fairness in employment decisions
New York City Local Law 144: Bias audits for automated employment decision tools
12.7 Compliance Mapping
This framework supports compliance across jurisdictions:
| Regulation | Framework Component | Evidence Provided |
| EU AI Act High-Risk | Tier 3 requirements | Risk management, TEVV, monitoring, transparency |
| EU AI Act GPAI | Tier 2-3 for foundation models | Technical docs, red-team testing, incident response |
| GDPR/UK GDPR | Privacy review, DPIA | Data handling, consent, DPIA in Risk Analysis |
| CCPA/CPRA | Privacy review, transparency | Consumer notices, opt-out mechanisms |
| US EO 14110 | Red-team testing, fairness | Security testing, bias audits |
| Sector-specific | Domain Lead, specialized TEVV | Clinical validation, model risk management |
13. Control Categories Overview
The framework includes 32 controls organized into 15 categories aligned with the NIST AI RMF lifecycle:
Category 1: Governance and Organization (GOVERN)
AAT-01: AI System Inventory
AAT-02: Risk Classification
AAT-03: Governance Framework
AAT-04: Role Assignment
AAT-05: Policy Documentation
Category 2: Data Management (MAP)
AAT-06: Training Data Documentation
AAT-07: Data Quality Standards
AAT-08: Bias Detection in Data
AAT-09: Data Minimization
Category 3: Model Development (MAP/MEASURE)
AAT-10: Model Development Standards
AAT-11: Performance Metrics Definition
AAT-12: Model Validation
AAT-13: Model Documentation
Category 4: Testing and Evaluation (MEASURE)
AAT-14: Pre-Deployment Testing
AAT-15: Security Testing
AAT-16: Adversarial Testing
AAT-17: User Acceptance Testing
Category 5: Deployment (MANAGE)
AAT-18: Deployment Controls
AAT-19: Access Controls
AAT-20: Monitoring Infrastructure
Category 6: Operations and Monitoring (MANAGE)
AAT-21: Performance Monitoring
AAT-22: Incident Response
AAT-23: Human Oversight
Category 7: Transparency and Explainability (All Functions)
AAT-24: User Transparency
AAT-25: Explainability
AAT-26: User Rights
Category 8: Maintenance and Evolution (MANAGE/MEASURE)
AAT-27: Change Management
AAT-28: Model Retraining
AAT-29: Version Control
Category 9: Decommissioning (MANAGE)
- AAT-29.1: Decommissioning Process
Category 10: Third-Party AI (GOVERN/MAP)
AAT-29.2: Third-Party AI Assessment
AAT-29.3: AI Supply Chain Due Diligence
AAT-29.4: Vendor Contract Terms
Category 11: Environmental and Social Impact (MAP)
AAT-29.5: Environmental Impact Assessment
AAT-29.6: Societal Impact Assessment
AAT-29.7: Stakeholder Engagement
Category 12: Accountability Mechanisms (GOVERN/MANAGE)
- AAT-29.8: Appeals/Redress Mechanism
Category 13: Compliance and Audit (GOVERN)
AAT-30: Audit Readiness
AAT-31: Record Retention
AAT-32: Compliance Reporting
Category 14: Ethics Review (GOVERN)
- AAT-32.1: AI Ethics Review
Category 15: Continuous Improvement (All Functions)
- Integrated across all controls through feedback loops
14. Complete Control Catalog (AAT-01 through AAT-16)
CATEGORY 1: GOVERNANCE AND ORGANIZATION (GOVERN)
AAT-01: AI System Inventory
Objective: Maintain a comprehensive, up-to-date inventory of all AI systems in use across the organization.
Control Question: Does the organization maintain a current inventory of all AI systems, including their risk classification, ownership, and status?
Implementation Guidance:
Create centralized registry of all AI systems
Include: System name, description, owner, tier, status, deployment date
Update within 5 business days of new system or status change
Review inventory quarterly for completeness
Make inventory accessible to governance team and auditors
Weighting by Tier:
Tier 1: Basic entry in inventory
Tier 2: Detailed inventory entry with all metadata
Tier 3: Enhanced tracking with executive visibility
PPTDF Mapping:
Pre-Processing: Inventory created at intake
Training: N/A
Deployment: Status updated upon deployment
Feedback: Regular inventory reviews
SCRM Considerations:
Track vendor-provided AI systems separately
Document supply chain dependencies
Monitor for vendor system changes
Related Standards:
ISO 42001: 6.1.2 (AI system inventory)
NIST AI RMF: GOVERN 1.2
EU AI Act: Article 71 (registration database)
AAT-02: Risk Classification
Objective: Systematically classify AI systems by risk level using a repeatable methodology.
Control Question: Are all AI systems classified using the risk tiering rubric, with justification documented?
Implementation Guidance:
Use 6-dimension scoring rubric (Section 8)
Document dimension scores and justifications
Apply automatic Tier 3 overrides where applicable
Re-tier upon material changes
AI Risk Lead validates tier assignment
Weighting by Tier:
Tier 1: Self-assessment acceptable
Tier 2: AI Risk Lead validation required
Tier 3: ARB confirms tier + automatic override check
PPTDF Mapping:
Pre-Processing: Initial tiering at intake
Training: Re-tier if training data changes significantly
Deployment: Validate tier before launch
Feedback: Re-tier based on incidents or reviews
SCRM Considerations:
Vendor system tier may differ from vendor's assessment
Consider supply chain attack surface in D6 (Misuse)
Related Standards:
ISO 42001: 6.1.3 (AI risk assessment)
NIST AI RMF: GOVERN 1.3, GOVERN 4.1
EU AI Act: Annex III (high-risk classification)
AAT-03: Governance Framework
Objective: Establish and maintain a formal AI governance framework defining policies, procedures, and oversight.
Control Question: Is there a documented AI governance framework that is reviewed annually and updated based on incidents or regulatory changes?
Implementation Guidance:
Publish governance standard (this document)
Define roles, responsibilities, decision rights
Establish AI Review Board (ARB)
Document approval workflows
Review annually or after major incidents/regulatory changes
Communicate updates to all stakeholders
Weighting by Tier:
Tier 1: Basic awareness of framework required
Tier 2: Full framework compliance
Tier 3: Enhanced governance with executive oversight
PPTDF Mapping:
Pre-Processing: Framework establishes requirements
Training: Framework guides training practices
Deployment: Framework defines approval gates
Feedback: Framework updated based on lessons learned
SCRM Considerations:
Framework must address third-party AI risks
Vendor governance requirements specified
Related Standards:
ISO 42001: 5.1 (AI policy), 5.2 (AI management system)
NIST AI RMF: GOVERN 1.1, GOVERN 2.1
EU AI Act: Article 9 (risk management system)
AAT-04: Role Assignment
Objective: Ensure clear accountability through defined roles for each AI system.
Control Question: Are required roles (Business Owner, Engineering Owner, etc.) assigned for each AI system with documented responsibilities?
Implementation Guidance:
Assign all required roles per tier (Section 7)
Document in Intake Form and System Card
Roles have authority and resources
CATEGORY 1: GOVERNANCE AND ORGANIZATION
AAT-04: Role Assignment (Continued)
Implementation Guidance (Continued):
Assign all required roles per tier (Section 7)
Document in Intake Form and System Card
Roles have authority and resources to fulfill responsibilities
Notify role holders of assignments and expectations
Track role changes and transitions
Weighting by Tier:
Tier 1: Business Owner + Engineering Owner
Tier 2: Tier 1 + Security + Privacy/Legal + Domain Lead
Tier 3: Tier 2 + Executive Risk Owner + Full ARB
PPTDF Mapping:
Pre-Processing: Roles assigned at intake
Training: Engineering/Domain Lead oversee training
Deployment: All roles approve before deployment
Feedback: Roles participate in post-launch reviews
SCRM Considerations:
Vendor relationship owner must be designated
Integration owner for third-party systems
Related Standards:
ISO 42001: 5.3 (Roles and responsibilities)
NIST AI RMF: GOVERN 1.5
EU AI Act: Article 16 (obligations of providers)
AAT-05: Policy Documentation
Objective: Document and communicate AI policies covering acceptable use, prohibited uses, and governance requirements.
Control Question: Are AI policies documented, approved, published, and reviewed annually?
Implementation Guidance:
Publish AI Governance Standard (this document)
Create Acceptable Use Policy for AI tools
Document prohibited AI use cases
Define data handling requirements
Communicate policies to all relevant personnel
Obtain acknowledgment of policies
Review and update annually
Weighting by Tier:
Tier 1: Awareness of policies required
Tier 2: Compliance verification required
Tier 3: Policy exceptions require executive approval
PPTDF Mapping:
Pre-Processing: Policies guide data collection
Training: Policies restrict training data use
Deployment: Policies enforced before launch
Feedback: Policy violations trigger reviews
SCRM Considerations:
Vendor policies must align with organizational policies
Contractual enforcement of policy compliance
Related Standards:
ISO 42001: 5.1 (AI policy)
NIST AI RMF: GOVERN 2.1
SOC 2: CC1.1 (Control environment)
CATEGORY 2: DATA MANAGEMENT (MAP)
AAT-06: Training Data Documentation
Objective: Document the sources, composition, and characteristics of training data.
Control Question: Is training data documented including sources, size, composition, collection methods, and known limitations?
Implementation Guidance:
Document all training data sources
Record data collection methodology
Describe data composition (domains, languages, demographics)
Identify data gaps and limitations
Document data provenance and lineage
Specify data licensing and usage rights
Update documentation when data changes
Weighting by Tier:
Tier 1: Basic description of data sources
Tier 2: Comprehensive data documentation
Tier 3: Detailed provenance with independent validation
PPTDF Mapping:
Pre-Processing: Document data sources and collection
Training: Record training data composition
Deployment: N/A
Feedback: Update based on data drift findings
SCRM Considerations:
Vendor must disclose training data sources
Third-party data licensing verified
Data supply chain risks assessed
Related Standards:
ISO 42001: 7.4 (Data for AI system)
NIST AI RMF: MAP 1.1
EU AI Act: Article 10 (data governance)
AAT-07: Data Quality Standards
Objective: Ensure training and operational data meet defined quality standards.
Control Question: Are data quality standards defined and validated, including accuracy, completeness, representativeness, and freshness?
Implementation Guidance:
Define quality metrics (accuracy, completeness, consistency, timeliness)
Establish quality thresholds
Validate data against quality standards
Document data cleaning and preprocessing
Monitor data quality over time
Address quality issues before training/deployment
Maintain data quality audit trail
Weighting by Tier:
Tier 1: Basic quality checks
Tier 2: Formal quality validation
Tier 3: Comprehensive quality assurance with independent audit
PPTDF Mapping:
Pre-Processing: Quality standards enforced during collection
Training: Quality validated before training
Deployment: Operational data quality monitored
Feedback: Quality metrics tracked continuously
SCRM Considerations:
Vendor data quality requirements specified
Data supplier quality audits required
Related Standards:
ISO 42001: 7.4.2 (Data quality)
NIST AI RMF: MAP 2.3
EU AI Act: Article 10.3 (data quality)
AAT-08: Bias Detection in Data
Objective: Identify and document bias in training and operational data.
Control Question: Has training data been assessed for bias, with findings documented and mitigation strategies implemented?
Implementation Guidance:
Analyze data for representational bias
Check demographic distribution vs. target population
Identify underrepresented groups
Document label bias and historical bias
Assess measurement bias
Implement bias mitigation strategies (re-sampling, re-weighting)
Document residual bias and limitations
Weighting by Tier:
Tier 1: Basic bias awareness
Tier 2: Formal bias assessment
Tier 3: Comprehensive bias audit with subgroup analysis
PPTDF Mapping:
Pre-Processing: Bias detection during data collection
Training: Bias mitigation techniques applied
Deployment: N/A
Feedback: Ongoing bias monitoring in production
SCRM Considerations:
Vendor must disclose bias assessments
Third-party data bias characteristics documented
Related Standards:
ISO 42001: 6.2 (Fairness)
NIST AI RMF: MAP 3.3, MEASURE 2.7
EU AI Act: Article 10.2(f) (relevant, representative data)
AAT-09: Data Minimization
Objective: Collect and process only data necessary for the AI system's purpose.
Control Question: Has data minimization been applied to limit collection to what is necessary, with justification for all data elements?
Implementation Guidance:
Define minimum data set needed for system purpose
Justify each data element collected
Remove unnecessary features from training data
Limit retention to required period
Implement privacy-preserving techniques where possible
Document data minimization decisions
Regular review of data necessity
Weighting by Tier:
Tier 1: Basic data minimization
Tier 2: Formal minimization analysis
Tier 3: Privacy-by-design with minimization audit
PPTDF Mapping:
Pre-Processing: Minimize data collection
Training: Feature selection reduces dimensionality
Deployment: Runtime data collection minimized
Feedback: Review data necessity regularly
SCRM Considerations:
Vendor data collection limited by contract
Data sharing minimized
Related Standards:
ISO 42001: 7.4.3 (Data minimization)
GDPR: Article 5.1(c) (Data minimization)
NIST AI RMF: MAP 1.5
CATEGORY 3: MODEL DEVELOPMENT (MAP/MEASURE)
AAT-10: Model Development Standards
Objective: Follow secure, reproducible model development practices.
Control Question: Are model development standards documented and followed, including version control, reproducibility, and secure coding?
Implementation Guidance:
Use version control for all code and configurations
Document model architecture decisions
Ensure reproducibility (fixed random seeds, versioned dependencies)
Follow secure coding practices
Implement code review process
Maintain development environment security
Document hyperparameter choices
Weighting by Tier:
Tier 1: Basic version control
Tier 2: Comprehensive development standards
Tier 3: Enhanced security and reproducibility requirements
PPTDF Mapping:
Pre-Processing: Standards apply to data preparation code
Training: Standards govern training scripts
Deployment: Deployment code follows standards
Feedback: Updates follow same standards
SCRM Considerations:
Vendor development practices assessed
Code review for integrated vendor models
Related Standards:
ISO 42001: 7.5 (AI system development)
NIST AI RMF: MAP 4.1
SOC 2: CC8.1 (Change management)
AAT-11: Performance Metrics Definition
Objective: Define clear, measurable performance metrics before development.
Control Question: Are performance metrics and acceptance thresholds defined and documented before model training?
Implementation Guidance:
Define task-specific metrics (accuracy, precision, recall, F1, etc.)
Set minimum acceptable thresholds
Define metrics for all trustworthiness characteristics
Document metric calculation methodology
Specify evaluation datasets
Define subgroup-specific metrics if applicable
Get stakeholder agreement on metrics
Weighting by Tier:
Tier 1: Basic performance metrics
Tier 2: Comprehensive metrics including fairness
Tier 3: Full trustworthiness metrics with subgroup analysis
PPTDF Mapping:
Pre-Processing: Metrics inform data requirements
Training: Metrics guide model selection
Deployment: Metrics set deployment thresholds
Feedback: Metrics monitored in production
SCRM Considerations:
Vendor must report on agreed metrics
SLAs tied to performance metrics
Related Standards:
ISO 42001: 7.5.2 (AI system performance)
NIST AI RMF: MEASURE 1.1
EU AI Act: Article 15 (accuracy requirements)
AAT-12: Model Validation
Objective: Validate model performance against defined metrics using independent test data.
Control Question: Has the model been validated on independent test data with results documented and compared to thresholds?
Implementation Guidance:
Use separate test set (not used in training)
Validate against all defined metrics
Test on diverse data representing deployment conditions
Document validation methodology
Compare results to acceptance thresholds
Investigate failures and edge cases
Independent validation for Tier 3
Weighting by Tier:
Tier 1: Basic validation on test set
Tier 2: Comprehensive validation including edge cases
Tier 3: Independent validation with detailed documentation
PPTDF Mapping:
Pre-Processing: N/A
Training: Validation follows training
Deployment: Validation gates deployment
Feedback: Validation repeated after retraining
SCRM Considerations:
Vendor validation results verified
Independent validation for vendor models (Tier 3)
Related Standards:
ISO 42001: 7.6 (Verification and validation)
NIST AI RMF: MEASURE 2.1
EU AI Act: Article 9.2(c) (testing procedures)
AAT-13: Model Documentation
Objective: Create comprehensive model documentation (model cards) for transparency and knowledge transfer.
Control Question: Is comprehensive model documentation maintained, including architecture, training process, performance, and limitations?
Implementation Guidance:
Complete model card (Appendix E template)
Document model architecture and hyperparameters
Describe training data and process
Report performance metrics
Document known limitations and failure modes
Specify intended use and prohibited uses
Update documentation with model changes
Weighting by Tier:
Tier 1: Basic model card
Tier 2: Comprehensive model card
Tier 3: Enhanced documentation with independent review
PPTDF Mapping:
Pre-Processing: Document data preparation
Training: Document training process
Deployment: Document deployment configuration
Feedback: Update based on production learnings
SCRM Considerations:
Vendor must provide model cards
Model card completeness verified
Related Standards:
ISO 42001: 7.7 (AI system documentation)
NIST AI RMF: MAP 5.1
EU AI Act: Article 11 (technical documentation)
CATEGORY 4: TESTING AND EVALUATION (MEASURE)
AAT-14: Pre-Deployment Testing
Objective: Conduct comprehensive TEVV testing before deployment.
Control Question: Has comprehensive pre-deployment testing been completed covering functionality, safety, reliability, fairness, privacy, and security?
Implementation Guidance:
Execute TEVV plan (Section 11.3 template)
Test functional correctness against requirements
Safety testing for harmful outputs (GenAI)
Reliability testing (regression, edge cases)
Fairness testing across subgroups (if applicable)
Privacy leakage testing
Security testing (prompt injection, etc.)
Document all test results
Retest after mitigations
Weighting by Tier:
Tier 1: Light TEVV (20-100 cases)
Tier 2: Full TEVV (100-500 cases)
Tier 3: Comprehensive TEVV (500+ cases) + independent validation
PPTDF Mapping:
Pre-Processing: N/A
Training: Testing follows training
Deployment: Testing gates deployment
Feedback: Testing repeated for updates
SCRM Considerations:
Vendor testing results verified
Additional testing for vendor integrations
Related Standards:
ISO 42001: 7.6 (Verification and validation)
NIST AI RMF: MEASURE 2.2
EU AI Act: Article 9.4 (testing)
AAT-15: Security Testing
Objective: Test AI system security controls and resistance to attacks.
Control Question: Has security testing been conducted including prompt injection, access control, and data exfiltration scenarios?
Implementation Guidance:
Threat model identifies test scenarios (Section 11.4)
Test prompt injection attacks (direct and indirect)
Test access control bypass attempts
Test data exfiltration via outputs
Test tool misuse (if agentic)
Test jailbreak attempts (GenAI)
Penetration testing (Tier 3)
Document findings and mitigations
Retest after fixes
Weighting by Tier:
Tier 1: Basic security checks
Tier 2: Comprehensive security testing
Tier 3: Red-team security testing
PPTDF Mapping:
Pre-Processing: N/A
Training: N/A
Deployment: Security testing before deployment
Feedback: Ongoing security monitoring
SCRM Considerations:
Vendor security testing verified
Integration security tested
Related Standards:
ISO 42001: 6.3 (Security)
NIST AI RMF: MEASURE 2.9
EU AI Act: Article 15 (cybersecurity)
OWASP Top 10 for LLMs
AAT-16: Adversarial Testing
Objective: Conduct red-team testing to identify vulnerabilities through adversarial attacks.
Control Question: Has adversarial/red-team testing been conducted by independent testers attempting to bypass controls and cause harm?
Implementation Guidance:
Engage independent red-team (internal or external)
Define attack scenarios (jailbreaks, social engineering, data extraction)
Document attack attempts and success rates
Identify successful attack patterns
Implement mitigations
Conduct retest to validate mitigations
Document residual vulnerabilities
Required for Tier 3; recommended for Tier 2
Weighting by Tier:
Tier 1: Not required
Tier 2: Recommended; basic adversarial testing
Tier 3: Required; comprehensive red-team testing
PPTDF Mapping:
Pre-Processing: N/A
Training: N/A
Deployment: Red-team testing before deployment
Feedback: Periodic red-team testing (annual for Tier 3)
SCRM Considerations:
Vendor systems require red-team testing
Integration points tested
Related Standards:
NIST AI 600-1: 2.10 (Adversarial testing for GenAI)
NIST AI RMF: MEASURE 2.10
EU AI Act: Article 9.4(d) (testing against specifications)
AAT-17: User Acceptance Testing
Objective: Validate system usability and fitness for purpose with representative users.
Control Question: Has user acceptance testing been conducted with representative users validating usability and fitness for intended use?
Implementation Guidance:
Identify representative user sample
Define UAT scenarios covering typical workflows
Collect user feedback on usability, accuracy, helpfulness
Document user concerns and confusion points
Assess user understanding of AI limitations
Validate user disclosure effectiveness
Address issues before launch
Required for Tier 2-3
Weighting by Tier:
Tier 1: Optional; informal user feedback
Tier 2: Formal UAT with representative users
Tier 3: Comprehensive UAT with diverse user groups
PPTDF Mapping:
Pre-Processing: N/A
Training: N/A
Deployment: UAT before deployment
Feedback: Ongoing user feedback collection
SCRM Considerations:
UAT includes vendor system integration
User feedback on vendor UX
Related Standards:
ISO 42001: 7.6.3 (Validation)
NIST AI RMF: MEASURE 2.3
ISO 9241 (Usability)
CATEGORY 5: DEPLOYMENT (MANAGE)
AAT-18: Deployment Controls
Objective: Implement controlled deployment with rollback capability.
Control Question: Are deployment controls in place including phased rollout, rollback capability, and kill switch (Tier 3)?
Implementation Guidance:
Deploy to staging environment first
Phased production rollout (canary → gradual → full)
Monitor closely during rollout
Maintain rollback capability to previous version
Implement kill switch for emergency shutdown (Tier 3)
Test rollback and kill switch before deployment
Document deployment steps and rollback procedure
Weighting by Tier:
Tier 1: Basic deployment with rollback
Tier 2: Phased rollout with monitoring
Tier 3: Phased rollout + tested kill switch
PPTDF Mapping:
Pre-Processing: N/A
Training: N/A
Deployment: Controls enforced during deployment
Feedback: Rollback used if issues detected
SCRM Considerations:
Vendor SLA includes rollback capabilities
Vendor service degradation triggers rollback
Related Standards:
ISO 42001: 8.2 (AI system deployment)
NIST AI RMF: MANAGE 1.3
SOC 2: CC8.1 (Change management)
AAT-19: Access Controls
Objective: Implement appropriate access controls for AI system usage and administration.
Control Question: Are access controls implemented using least privilege principles with regular access reviews?
Implementation Guidance:
Implement role-based access control (RBAC)
Apply least privilege principle
Separate user access from admin access
Control access to prompts, outputs, logs, model weights
Implement authentication and authorization
Monitor access attempts and anomalies
Regular access reviews (quarterly for Tier 2-3)
Revoke access promptly when no longer needed
Weighting by Tier:
Tier 1: Basic access controls
Tier 2: RBAC with regular reviews
Tier 3: Enhanced access controls with monitoring and auditing
PPTDF Mapping:
Pre-Processing: Access to training data controlled
Training: Access to training infrastructure limited
Deployment: Runtime access controls enforced
Feedback: Access logs reviewed regularly
SCRM Considerations:
Vendor access controls verified
Vendor admin access monitored
Related Standards:
ISO 42001: 6.3.2 (Access control)
NIST AI RMF: GOVERN 3.2
SOC 2: CC6.1 (Logical access)
GDPR: Article 32 (Security)
AAT-20: Monitoring Infrastructure
Objective: Implement comprehensive monitoring infrastructure before deployment.
Control Question: Is monitoring infrastructure operational before deployment, capturing performance, safety, fairness, and security metrics?
Implementation Guidance:
Define monitoring metrics based on TEVV
Implement monitoring dashboards
Configure alerting thresholds
Set up logging infrastructure
Ensure monitoring operational before launch
Test alerting functionality
Document monitoring and alert response procedures
Assign on-call responsibilities
Weighting by Tier:
Tier 1: Basic usage and error monitoring
Tier 2: Comprehensive monitoring with alerting
Tier 3: Real-time monitoring with executive dashboards
PPTDF Mapping:
Pre-Processing: N/A
Training: N/A
Deployment: Monitoring required before deployment
Feedback: Monitoring provides continuous feedback
SCRM Considerations:
Vendor monitoring capabilities verified
Vendor SLA metrics tracked
Related Standards:
ISO 42001: 8.3 (Use of AI system)
NIST AI RMF: MANAGE 4.1
SOC 2: CC7.1 (System monitoring)
CATEGORY 6: OPERATIONS AND MONITORING (MANAGE)
AAT-21: Performance Monitoring
Objective: Continuously monitor AI system performance against defined SLOs.
Control Question: Is system performance continuously monitored against SLOs with alerts for degradation or drift?
Implementation Guidance:
Monitor key performance metrics from TEVV
Track accuracy, latency, availability
Monitor for data drift and model drift
Monitor for concept drift
Alert on SLO violations
Investigate performance degradation
Escalate persistent issues
Document monitoring findings
Weighting by Tier:
Tier 1: Weekly performance reviews
Tier 2: Daily automated monitoring with alerts
Tier 3: Real-time monitoring with immediate alerts
PPTDF Mapping:
Pre-Processing: N/A
Training: N/A
Deployment: N/A
Feedback: Continuous performance monitoring provides feedback
SCRM Considerations:
Vendor performance monitored per SLA
Vendor performance issues escalated
Related Standards:
ISO 42001: 8.3.2 (Monitoring)
NIST AI RMF: MANAGE 4.2
EU AI Act: Article 72 (Post-market monitoring)
AAT-22: Incident Response
Objective: Respond effectively to AI system incidents with defined procedures.
Control Question: Is there a documented AI incident response plan with defined severity levels, escalation paths, and response procedures?
Implementation Guidance:
Implement AI Incident Response Playbook (Appendix F)
Define incident severity levels (P1-P4)
Establish incident response team
Document escalation procedures
Train team on incident response
Conduct incident response drills (Tier 3)
Post-incident reviews and documentation
Update playbook based on lessons learned
Weighting by Tier:
Tier 1: Basic incident handling
Tier 2: Documented incident response procedures
Tier 3: Comprehensive playbook with drills and executive escalation
PPTDF Mapping:
Pre-Processing: N/A
Training: N/A
Deployment: N/A
Feedback: Incident learnings drive improvements
SCRM Considerations:
Vendor incident notification requirements
Vendor incident response coordination
Related Standards:
ISO 42001: 8.3.3 (Incident management)
NIST AI RMF: MANAGE 2.2
EU AI Act: Article 73 (Serious incidents reporting)
SOC 2: CC7.3 (Incident response)
AAT-23: Human Oversight
Objective: Implement appropriate human oversight mechanisms for AI decisions and outputs.
Control Question: Are human oversight mechanisms implemented appropriate to the system's autonomy level and risk?
Implementation Guidance:
Define human review points in workflow
Implement human-in-the-loop for high-stakes decisions
Enable human override capability
Provide adequate context for human review
Train reviewers on AI limitations
Monitor human oversight effectiveness
Document override rationale
Review override patterns for system improvement
Weighting by Tier:
Tier 1: User reviews outputs before use
Tier 2: Defined review points with documentation
Tier 3: Mandatory human review for high-stakes decisions
PPTDF Mapping:
Pre-Processing: N/A
Training: N/A
Deployment: Oversight mechanisms operational at deployment
Feedback: Override patterns inform improvements
SCRM Considerations:
Vendor systems integrated with oversight workflows
Vendor outputs subject to review requirements
Related Standards:
ISO 42001: 6.4 (Human oversight)
NIST AI RMF: MANAGE 1.2
EU AI Act: Article 14 (Human oversight)
GDPR: Article 22 (Automated decision-making)
CATEGORY 7: TRANSPARENCY AND EXPLAINABILITY
AAT-24: User Transparency
Objective: Ensure users are informed when interacting with AI systems.
Control Question: Are users clearly informed when they are interacting with an AI system, with appropriate disclosures about capabilities and limitations?
Implementation Guidance:
Display clear AI disclosure to users
Explain AI system purpose and capabilities
Communicate known limitations
Provide guidance on appropriate use
Disclose when content is AI-generated (GenAI)
Make disclosures prominent and understandable
Test user comprehension of disclosures
Weighting by Tier:
Tier 1: Basic AI disclosure
Tier 2: Comprehensive disclosure with limitations
Tier 3: Detailed transparency with provenance (GenAI)
PPTDF Mapping:
Pre-Processing: N/A
Training: N/A
Deployment: Disclosures presented to users
Feedback: User questions about AI inform disclosure updates
SCRM Considerations:
Vendor AI systems properly disclosed
Vendor branding vs. organizational branding clarified
Related Standards:
ISO 42001: 6.5 (Transparency)
NIST AI RMF: GOVERN 5.1
EU AI Act: Article 50 (Transparency obligations)
CCPA: Automated decision-making disclosure
AAT-25: Explainability
Objective: Provide explanations for AI system outputs appropriate to user needs and system risk.
Control Question: Are explanations available for AI outputs appropriate to the system's risk level and user needs?
Implementation Guidance:
Implement explainability techniques (LIME, SHAP, attention, etc.)
Provide user-facing explanations when needed
Explain decision factors for high-stakes decisions
Document explainability limitations
Test explanation quality and user comprehension
Balance explainability with other objectives (accuracy, privacy)
Weighting by Tier:
Tier 1: Optional; basic explanations
Tier 2: Explanations for key decisions
Tier 3: Comprehensive explainability for high-stakes decisions
PPTDF Mapping:
Pre-Processing: N/A
Training: Explainability considerations in model selection
Deployment: Explanations available at runtime
Feedback: User questions about decisions inform explanations
SCRM Considerations:
Vendor explainability capabilities assessed
Black-box vendor models may limit explainability
Related Standards:
ISO 42001: 6.5.2 (Explainability)
NIST AI RMF: MEASURE 2.8
EU AI Act: Recital 47 (Explainability)
GDPR: Recital 71 (Right to explanation)
AAT-26: User Rights
Objective: Enable users to exercise rights regarding AI-based decisions.
Control Question: Are user rights mechanisms implemented including appeal, correction, and opt-out where applicable?
Implementation Guidance:
Provide mechanism to challenge AI decisions
Enable users to request human review
Allow users to correct inaccurate data
Implement opt-out for certain AI processing (where required)
Document and respond to user rights requests
Track and analyze rights requests
Required for Tier 2-3
Weighting by Tier:
Tier 1: Basic user feedback mechanism
Tier 2: Formal rights request process
Tier 3: Comprehensive rights with appeal/redress mechanism
PPTDF Mapping:
Pre-Processing: N/A
Training: N/A
Deployment: Rights mechanisms operational
Feedback: Rights requests inform system improvements
SCRM Considerations:
Vendor systems support user rights requests
Vendor cooperation in responding to requests
Related Standards:
GDPR: Articles 15-22 (Data subject rights)
CCPA: Sections 1798.100-1798.125 (Consumer rights)
EU AI Act: Article 86 (Right to explanation)
CATEGORY 8: MAINTENANCE AND EVOLUTION (MANAGE/MEASURE)
AAT-27: Change Management
Objective: Manage changes to AI systems through controlled processes.
Control Question: Is there a documented change management process that determines when changes require re-approval and re-testing?
Implementation Guidance:
Define material vs. non-material changes
Material changes require re-tiering and re-approval
Implement change request process
Document change impact assessment
Test changes before deployment
Maintain change log
Rollback capability for changes
Weighting by Tier:
Tier 1: Basic change tracking
Tier 2: Formal change management with approval
Tier 3: Rigorous change control with executive approval for material changes
PPTDF Mapping:
Pre-Processing: Changes to data require review
Training: Model changes trigger change management
Deployment: Deployment changes controlled
Feedback: Change impact monitored post-deployment
SCRM Considerations:
Vendor changes communicated and assessed
Vendor change notice requirements in contract
Related Standards:
ISO 42001: 8.4 (Change management)
NIST AI RMF: MANAGE 3.1
SOC 2: CC8.1 (Change management)
AAT-28: Model Retraining
Objective: Retrain models systematically to maintain performance and address drift.
Control Question: Is there a defined retraining schedule and process with triggers for unscheduled retraining based on drift detection?
Implementation Guidance:
Define retraining frequency (monthly, quarterly, etc.)
Monitor for drift triggers
Document retraining process and data
Validate retrained model before deployment
Compare new model to previous version (regression testing)
Update model documentation
Deploy retrained model via change management
Weighting by Tier:
Tier 1: Retraining as needed
Tier 2: Scheduled retraining with drift monitoring
Tier 3: Frequent retraining with comprehensive validation
PPTDF Mapping:
Pre-Processing: New training data collected
Training: Retraining process executed
Deployment: Retrained model deployed via change management
Feedback: Drift detection triggers retraining
SCRM Considerations:
Vendor retraining schedules disclosed
Vendor model updates assessed for impact
Related Standards:
ISO 42001: 8.3.4 (Continual learning)
NIST AI RMF: MANAGE 4.3
EU AI Act: Article 72.2 (Updating of datasets)
AAT-29: Version Control
Objective: Maintain version control for all AI system components.
Control Question: Are all AI system components version-controlled including code, models, configurations, and data?
Implementation Guidance:
Use version control system (Git, etc.)
Version all code, configurations, prompts
Version model weights and artifacts
Version training data (lineage tracking)
Tag releases with version numbers
Maintain version history and changelog
Enable rollback to previous versions
Weighting by Tier:
Tier 1: Basic code version control
Tier 2: Comprehensive version control including models
Tier 3: Full versioning with immutable audit trail
PPTDF Mapping:
Pre-Processing: Data versions tracked
Training: Training runs versioned
Deployment: Deployed versions documented
Feedback: Version correlation with issues tracked
SCRM Considerations:
Vendor model versions documented
Vendor version change notifications required
Related Standards:
ISO 42001: 7.5.4 (Version control)
NIST AI RMF: MANAGE 3.2
SOC 2: CC8.1 (Change management)
CATEGORY 9: DECOMMISSIONING (MANAGE)
AAT-29.1: Decommissioning Process
Objective: Safely decommission AI systems with proper data handling and documentation.
Control Question: Is there a documented decommissioning process that includes data deletion, user notification, and knowledge transfer?
Implementation Guidance:
Create decommissioning plan before initiating shutdown
Notify users of decommissioning timeline
Export necessary data for compliance/legal retention
Delete unnecessary data per retention policy
Revoke system access and credentials
Archive system documentation
Document lessons learned
Transfer knowledge to successor system (if applicable)
Update system inventory to "decommissioned" status
Weighting by Tier:
Tier 1: Basic data deletion and notification
Tier 2: Formal decommissioning plan with documentation
Tier 3: Comprehensive decommissioning with audit trail
PPTDF Mapping:
Pre-Processing: Historical data handled per policy
Training: Model artifacts archived or deleted
Deployment: Infrastructure deprovisioned
Feedback: Decommissioning learnings documented
SCRM Considerations:
Vendor data deletion verified
Vendor contract termination procedures followed
Related Standards:
ISO 42001: 8.5 (Discontinuation)
NIST AI RMF: MANAGE 1.4
GDPR: Article 17 (Right to erasure)
CATEGORY 10: THIRD-PARTY AI (GOVERN/MAP)
AAT-29.2: Third-Party AI Assessment
Objective: Assess third-party AI systems before procurement and integration.
Control Question: Are third-party AI systems assessed using the vendor questionnaire with findings documented before procurement?
Implementation Guidance:
Use Vendor Assessment Questionnaire (Appendix H)
Assess vendor governance, security, privacy, fairness practices
Review vendor certifications (SOC 2, ISO 27001, etc.)
Evaluate vendor AI transparency and documentation
Score vendor across assessment dimensions
Document assessment findings and risks
Approve or reject based on risk assessment
Required for all new AI vendors
Weighting by Tier:
Tier 1: Basic vendor assessment
Tier 2: Comprehensive vendor assessment
Tier 3: Enhanced vendor assessment with on-site review
PPTDF Mapping:
Pre-Processing: Vendor data practices assessed
Training: Vendor training practices evaluated
Deployment: Vendor deployment controls reviewed
Feedback: Vendor monitoring capabilities assessed
SCRM Considerations:
Core control for SCRM
Vendor assessment repeated annually
Related Standards:
ISO 42001: 7.3 (Outsourced processes)
NIST AI RMF: GOVERN 6.1
EU AI Act: Article 16.2 (Provider due diligence)
AAT-29.3: AI Supply Chain Due Diligence
Objective: Conduct due diligence on the AI supply chain including model providers, data sources, and infrastructure.
Control Question: Has due diligence been conducted on all AI supply chain components including subprocessors, data sources, and infrastructure providers?
Implementation Guidance:
Map complete AI supply chain
Identify all third-party dependencies
Assess each supplier's risk profile
Review subprocessor agreements and DPAs
Verify data source licenses and rights
Assess infrastructure security (cloud providers)
Document supply chain risks
Implement mitigation strategies
Monitor supply chain continuously
Weighting by Tier:
Tier 1: Basic supply chain mapping
Tier 2: Comprehensive due diligence
Tier 3: Enhanced due diligence with audit rights
PPTDF Mapping:
Pre-Processing: Data source due diligence
Training: Training infrastructure assessed
Deployment: Deployment infrastructure verified
Feedback: Supply chain monitored continuously
SCRM Considerations:
Core control for SCRM
Fourth-party risk assessed (vendor's vendors)
Related Standards:
ISO 42001: 8.1 (Supply chain)
NIST AI RMF: GOVERN 6.2
EU AI Act: Article 16 (Obligations along the AI value chain)
AAT-29.4: Vendor Contract Terms
Objective: Ensure vendor contracts include appropriate AI-specific terms and protections.
Control Question: Do vendor contracts include required AI governance terms covering data usage, performance SLAs, security, indemnification, and termination?
Implementation Guidance:
Include data processing agreement (DPA)
Specify data usage restrictions (no training on customer data without consent)
Define performance SLAs with metrics
Require security and privacy certifications
Include audit rights
Specify incident notification timelines
Define indemnification for IP, data breaches, discrimination
Include termination and data return provisions
Require compliance with applicable AI regulations
Privacy/Legal review all AI vendor contracts
Weighting by Tier:
Tier 1: Standard DPA and data terms
Tier 2: Comprehensive AI contract terms
Tier 3: Enhanced terms with audit rights and indemnification
PPTDF Mapping:
Pre-Processing: Data terms govern data collection
Training: Training restrictions specified
Deployment: SLAs define deployment requirements
Feedback: Contract enables ongoing governance
SCRM Considerations:
Core control for SCRM
Contract terms enable vendor oversight
Related Standards:
ISO 42001: 7.3.2 (Agreements)
GDPR: Article 28 (Processor agreements)
NIST AI RMF: GOVERN 6.1
CATEGORY 11: ENVIRONMENTAL AND SOCIAL IMPACT (MAP)
AAT-29.5: Environmental Impact Assessment
Objective: Assess and document the environmental impact of AI systems, particularly energy consumption.
Control Question: Has the environmental impact of the AI system been assessed including training and inference energy consumption?
Implementation Guidance:
Estimate training compute and energy consumption
Calculate carbon footprint (CO2 equivalent)
Estimate inference energy per request
Consider data center efficiency and energy sources
Implement efficiency optimizations where feasible
Document environmental impact in System Card
Report environmental metrics (Tier 3)
Recommended for all tiers; required for Tier 3 high-compute systems
Weighting by Tier:
Tier 1: Optional assessment
Tier 2: Basic environmental assessment
Tier 3: Comprehensive assessment with reporting
PPTDF Mapping:
Pre-Processing: Data processing energy considered
Training: Training energy calculated
Deployment: Inference efficiency optimized
Feedback: Energy consumption monitored
SCRM Considerations:
Vendor environmental impact disclosed
Cloud provider renewable energy usage
Related Standards:
ISO 42001: 6.6 (Sustainability)
EU AI Act: Recital 23 (Energy efficiency)
AAT-29.6: Societal Impact Assessment
Objective: Assess broader societal impacts of AI systems beyond individual users.
Control Question: Has a societal impact assessment been conducted considering effects on communities, labor markets, and vulnerable populations?
Implementation Guidance:
Identify stakeholders beyond direct users
Assess impact on vulnerable populations
Consider labor market effects (job displacement)
Evaluate potential for dual use or misuse
Consider environmental justice impacts
Assess effects on public discourse or democracy (if applicable)
Document mitigation strategies
Required for Tier 3; recommended for Tier 2
Weighting by Tier:
Tier 1: Not required
Tier 2: Basic societal impact consideration
Tier 3: Comprehensive societal impact assessment
PPTDF Mapping:
Pre-Processing: N/A
Training: N/A
Deployment: Impact assessment informs deployment decisions
Feedback: Ongoing impact monitoring
SCRM Considerations:
Vendor societal impact assessed
Combined organizational + vendor impact considered
Related Standards:
ISO 42001: 6.7 (Societal and environmental wellbeing)
NIST AI RMF: MAP 1.6
EU AI Act: Recital 27 (Societal impact)
AAT-29.7: Stakeholder Engagement
Objective: Engage relevant stakeholders in AI system development and governance.
Control Question: Have relevant stakeholders been identified and engaged in the development, deployment, and governance of the AI system?
Implementation Guidance:
Identify stakeholder groups (users, affected parties, domain experts, regulators)
Engage stakeholders in requirements and design
Solicit feedback on AI system impacts
Incorporate stakeholder concerns in risk assessment
Communicate AI capabilities and limitations
Establish feedback mechanisms
Document stakeholder engagement activities
Recommended for Tier 2; required for Tier 3
Weighting by Tier:
Tier 1: Basic user feedback
Tier 2: Structured stakeholder engagement
Tier 3: Comprehensive stakeholder engagement with documentation
PPTDF Mapping:
Pre-Processing: Stakeholders inform data requirements
Training: Stakeholders provide domain expertise
Deployment: Stakeholders validate fitness for purpose
Feedback: Ongoing stakeholder feedback
SCRM Considerations:
Stakeholder engagement includes vendor representatives
Affected parties may include vendor's customers
Related Standards:
ISO 42001: 4.2 (Understanding needs and expectations of interested parties)
NIST AI RMF: GOVERN 5.2
CATEGORY 12: ACCOUNTABILITY MECHANISMS (GOVERN/MANAGE)
AAT-29.8: Appeals/Redress Mechanism
Objective: Provide mechanisms for individuals to appeal or seek redress for AI decisions.
Control Question: Is there a documented appeals process for individuals affected by AI decisions, with clear procedures and timelines?
Implementation Guidance:
Establish appeals submission mechanism
Define appeal review process
Specify review timeline (e.g., 30 days)
Enable human review of appealed decisions
Document appeal decisions and rationale
Track appeal patterns and outcomes
Use appeals to improve system
Required for Tier 3 high-stakes decisions; recommended for Tier 2
Weighting by Tier:
Tier 1: Not required
Tier 2: Basic appeal mechanism
Tier 3: Formal appeals process with documentation and oversight
PPTDF Mapping:
Pre-Processing: N/A
Training: N/A
Deployment: Appeals mechanism operational
Feedback: Appeals inform system improvements
SCRM Considerations:
Appeals process covers vendor system decisions
Vendor cooperation in appeal reviews
Related Standards:
EU AI Act: Article 86 (Right to lodge a complaint)
GDPR: Article 77 (Right to lodge a complaint)
NIST AI RMF: GOVERN 5.3
CATEGORY 13: COMPLIANCE AND AUDIT (GOVERN)
AAT-30: Audit Readiness
Objective: Maintain audit readiness with organized documentation and evidence.
Control Question: Are all required governance artifacts maintained, organized, and accessible for internal and external audits?
Implementation Guidance:
Organize all governance artifacts by system
Maintain evidence of approvals and reviews
Document audit trail of decisions
Store artifacts in secure, accessible repository
Implement document retention policy
Conduct internal audits (annual for Tier 2-3)
Address audit findings promptly
Train team on audit procedures
Weighting by Tier:
Tier 1: Basic documentation retention
Tier 2: Organized audit trail
Tier 3: Comprehensive audit readiness with regular internal audits
PPTDF Mapping:
Pre-Processing: Documentation from start
Training: Training artifacts retained
Deployment: Deployment evidence maintained
Feedback: Ongoing documentation updated
SCRM Considerations:
Vendor audit reports obtained (SOC 2, etc.)
Vendor audit cooperation rights in contract
Related Standards:
ISO 42001: 9.2 (Internal audit)
SOC 2: CC4.1 (Monitoring activities)
EU AI Act: Article 64 (Access to documentation)
AAT-31: Record Retention
Objective: Retain AI system records per legal and regulatory requirements.
Control Question: Are record retention requirements defined and enforced, with records retained for the required period?
Implementation Guidance:
Define retention periods by record type
Consider legal, regulatory, and business requirements
Retain governance artifacts (System Cards, TEVV reports, etc.)
Retain model artifacts and training data (as feasible)
Retain operational logs per policy
Implement secure archival storage
Automate deletion after retention period
Document retention policy
Retention Recommendations:
Governance artifacts: 7 years minimum (Tier 3)
Model artifacts: Duration of deployment + 3 years
Training data: As required by regulation or until decommissioning
Operational logs: 90 days to 2 years depending on sensitivity
Weighting by Tier:
Tier 1: 1-3 years retention
Tier 2: 3-5 years retention
Tier 3: 7+ years retention
PPTDF Mapping:
Pre-Processing: Data lineage retained
Training: Training records archived
Deployment: Deployment logs retained
Feedback: Incident reports archived
SCRM Considerations:
Vendor retention requirements specified
Vendor data deletion verified
Related Standards:
ISO 42001: 7.7.2 (Record retention)
GDPR: Article 5.1(e) (Storage limitation)
SOX: Section 802 (Record retention)
EU AI Act: Article 12 (Record-keeping)
AAT-32: Compliance Reporting
Objective: Report on AI governance compliance to leadership, regulators, and stakeholders.
Control Question: Are compliance reports generated regularly and distributed to appropriate stakeholders including leadership and regulators?
Implementation Guidance:
Generate compliance reports (monthly/quarterly)
Report on governance program metrics
Track system coverage, approval velocity, incidents
Report to executive leadership (quarterly minimum)
Report to Board for high-risk systems (Tier 3)
Regulatory reporting as required (EU AI Act, sector-specific)
Document compliance gaps and remediation
Trending and improvement recommendations
Weighting by Tier:
Tier 1: Included in aggregate reporting
Tier 2: Quarterly reporting to leadership
Tier 3: Monthly reporting + Board reporting + regulatory reporting
PPTDF Mapping:
Pre-Processing: N/A
Training: N/A
Deployment: N/A
Feedback: Compliance reporting drives improvements
SCRM Considerations:
Vendor compliance reports obtained
Third-party risk included in compliance reporting
Related Standards:
ISO 42001: 9.3 (Management review)
SOC 2: CC4.2 (Communication of findings)
EU AI Act: Article 49 (Reporting of serious incidents)
CATEGORY 14: ETHICS REVIEW (GOVERN)
AAT-32.1: AI Ethics Review
Objective: Conduct ethics review for AI systems raising ethical considerations.
Control Question: Have AI systems with ethical implications been reviewed by an ethics board or committee with diverse perspectives?
Implementation Guidance:
Establish AI Ethics Review Board or Committee
Include diverse perspectives (technical, legal, social science, affected communities)
Define triggers for ethics review (high-stakes decisions, vulnerable populations, novel uses)
Conduct ethics review for Tier 3 systems (required)
Document ethical considerations and decisions
Ethics review informs risk assessment and design
Publish ethics principles and decisions (where appropriate)
Review ethics framework annually
Weighting by Tier:
Tier 1: Not required
Tier 2: Optional; recommended for ethically sensitive systems
Tier 3: Required ethics review
PPTDF Mapping:
Pre-Processing: Ethics informs data collection practices
Training: Ethics guides model development choices
Deployment: Ethics review gates deployment
Feedback: Ethical issues trigger review
SCRM Considerations:
Vendor ethics practices assessed
Vendor systems subject to organizational ethics review
Related Standards:
ISO 42001: 4.1 (Understanding the organization and its context - including ethical considerations)
NIST AI RMF: GOVERN 1.4
IEEE 7000 series (Ethics in design)
15. Implementation Guidance
15.1 Four-Phase Implementation Approach
Organizations should implement this framework in four phases:
Phase 1: Foundation (Months 1-3)
Objectives:
Establish governance structure
Create initial policies and templates
Begin system inventory
Key Activities:
Secure executive sponsorship and resources
Establish AI Review Board (ARB)
Assign AI Risk Lead role
Publish AI Governance Standard (this document)
Customize templates for organization
Begin AI system inventory
Communicate governance requirements to teams
Deliverables:
Published AI Governance Standard
Established ARB with charter
Customized templates
Initial system inventory (Tier 2-3 systems minimum)
Communication and training plan
Success Metrics:
Executive approval obtained
ARB operational with first meeting held
100% of known Tier 3 systems inventoried
Governance standard published and accessible
Phase 2: Pilot (Months 4-6)
Objectives:
Pilot governance process with 3-5 systems
Refine templates and workflows
Build organizational capability
Key Activities:
Select pilot systems (mix of tiers)
Apply governance process to pilots
Collect feedback from teams
Refine templates and guidance
Develop training materials
Begin training rollout (Level 1 and 2)
Establish governance tools/systems
Deliverables:
3-5 systems approved through governance process
Refined templates based on feedback
Training materials (Levels 1-2)
Initial training completion (50% target)
Governance tracking system operational
Success Metrics:
All pilot systems successfully approved
Average time-to-approval meets SLAs
80%+ team satisfaction with process
50% of AI personnel trained (Levels 1-2)
Phase 3: Scale (Months 7-12)
Objectives:
Scale governance across organization
Achieve comprehensive system coverage
Embed governance in culture
Key Activities:
Require governance for all new AI systems
Backfill governance for existing systems (prioritize Tier 2-3)
Complete training rollout
Implement monitoring and reporting
Conduct first internal audits
Refine processes based on experience
Develop advanced training (Levels 3-4)
Deliverables:
100% of Tier 3 systems governed
80%+ of Tier 2 systems governed
100% training completion (all required personnel)
Monthly compliance reports
First internal audit completed
Success Metrics:
100% Tier 3 coverage, 80% Tier 2 coverage
Average time-to-approval within SLA
Zero Tier 3 launches without approval
90%+ training completion
Clean internal audit (no critical findings)
Phase 4: Optimize (Ongoing)
Objectives:
Continuous improvement
Maintain compliance
Adapt to regulatory changes
Key Activities:
Achieve and maintain 100% system coverage
Quarterly program reviews
Annual framework review and updates
Benchmark against industry practices
Regular training updates
Proactive regulatory monitoring
Advanced capability development (automation, AI for governance)
Deliverables:
Quarterly compliance reports
Annual governance framework update
Updated training materials
Benchmark analysis
Continuous improvement roadmap
Success Metrics:
100% system coverage maintained
Time-to-approval continuously improving
Zero significant compliance incidents
Industry-leading governance maturity
Positive audit and regulatory feedback
15.2 Quick Start Checklist
Week 1:
[ ] Secure executive sponsorship
[ ] Assign AI Risk Lead
[ ] Form AI Review Board
[ ] Communicate governance initiative
Week 2-4:
[ ] Customize governance framework for organization
[ ] Customize templates (Section 11)
[ ] Create governance tracking system
[ ] Begin system inventory
Week 5-8:
[ ] Publish governance standard
[ ] Launch training program (Levels 1-2)
[ ] Select pilot systems
[ ] Hold first ARB meeting
Week 9-12:
[ ] Complete pilot system reviews
[ ] Refine templates based on feedback
[ ] Expand to all new systems
[ ] Begin backfilling existing systems
Month 4+:
[ ] Scale across organization
[ ] Monitor compliance metrics
[ ] Conduct internal audits
[ ] Continuous improvement
16. Appendix A: Filled Examples (Tier 1/2/3)
A.1 Tier 1 Example: Internal Meeting Summarizer
Use Case: Summarize meeting notes into action items for a single internal team (Sales team).
Risk Tier Score: 3
D1 Exposure: 0 (Internal team only)
D2 Impact: 0 (Convenience only)
D3 Autonomy: 1 (Generates draft; user reviews)
D4 Data: 1 (Internal non-sensitive meetings)
D5 Model/Change: 1 (Stable model, monthly updates)
D6 Misuse: 0 (Low abuse value)
Final Tier: 1
Approvals: Business Owner (VP Sales) + Engineering Owner
System Card Highlights:
Purpose: Internal-only summarization of sales team meetings; not for publishing or decisions
Model: OpenAI GPT-4o via API
Data Handling: No raw notes stored; metadata only logged (meeting ID, timestamp, user ID); 30-day retention
Human Oversight: User reviews output before copying; UI warning: "AI-generated summary - verify accuracy"
Monitoring: Thumbs up/down feedback + issue reports; kill switch via feature flag
Limitations: May miss nuances; not suitable for legal or compliance meetings
Light TEVV Summary:
60 sanitized sample meetings
Rubric: Faithfulness to original, completeness, action-item quality
Threshold: Invented action items ≤ 5%
Initial Results: 3.3% invented items
Mitigation: Added constraint "only extract explicit action items"
Retest Results: 1.6% invented items (PASS)
Risk Analysis Excerpt:
Top Risk #1: Hallucinated action items lead to missed commitments
Mitigation: UI warning + user review requirement
Residual: Low (user catches errors)
Top Risk #2: Accidental paste of sensitive information
Mitigation: No raw logging + access controls + team training
Residual: Low
Post-Launch: 45-day review confirmed 98% user satisfaction; 0 incidents; tier remains appropriate.
A.2 Tier 2 Example: Internal HR Policy Assistant
Use Case: HR staff ask questions; assistant drafts responses citing internal policy sources (RAG).
Risk Tier Score: 9
D1 Exposure: 1 (Organization-wide internal)
D2 Impact: 2 (Guidance affects employee decisions)
D3 Autonomy: 1 (Draft generation; HR staff reviews)
D4 Data: 2 (Employee questions may contain PII; policy data)
D5 Model/Change: 2 (LLM with RAG; weekly model updates)
D6 Misuse: 1 (Wrong guidance risk)
Final Tier: 2
Approvals: ARB Delegate + Security + Privacy/Legal + Domain Lead (CHRO)
Threat Model Highlights:
Scenario 1: Prompt Injection to Reveal Restricted Content
Attack: User crafts prompt to access confidential policies
Impact: Exposure of sensitive HR information
Controls: Allowlisted retrieval sources + snippet limits + output filters + access logs
Residual: Low
Scenario 2: PII Leakage into Logs
Attack: Employee names/details logged inadvertently
Impact: Privacy violation
Controls: Pre-log PII redaction + access controls + 90-day retention + encryption
Residual: Low (verified in testing)
Scenario 3: Hallucinated Policy Guidance
Attack: System generates plausible but incorrect guidance
Impact: Wrong advice to employees
Controls: Cite-or-refuse policy + source verification + hallucination detection
Residual: Medium (accepted with monitoring)
TEVV Report Highlights:
150-question test set covering common HR queries
Citation Coverage Target: ≥ 95%
Hallucination Target: ≤ 3%
Pre-Mitigation Results:
Citation coverage: 92% (FAIL)
Hallucination rate: 4.6% (FAIL)
Mitigations Implemented:
Added post-generation fact-checker
Implemented cite-first prompt engineering
Configured system to refuse if no citation available
Post-Mitigation Results:
Citation coverage: 97.2% (PASS)
Hallucination rate: 2.2% (PASS)
Privacy Audit: PII redaction tested; 0 leaks in 1000-sample audit (PASS)
Risk Analysis - Risk Register Excerpt:
Risk-001: Hallucinated Policy Guidance
Inherent Rating: 12 (High severity × Medium likelihood)
Controls: Citation requirement + hallucination detection + user warnings
Residual Rating: 6 (Medium severity × Low likelihood)
Acceptance Criteria: Hallucination ≤ 3%, citation ≥ 95%, monitored monthly
Owner: Engineering Owner
Status: ACCEPTED (within thresholds)
Risk-002: PII in Logs
Inherent Rating: 10 (High severity × Medium likelihood)
Controls: Pre-log PII redaction + access controls + encryption + 90-day retention
Residual Rating: 3 (Low severity × Low likelihood)
Acceptance Criteria: 0 redaction failures in monthly audits
Owner: Privacy Officer
Status: ACCEPTED (controls verified)
Monitoring:
Hallucination rate tracked monthly (alert if >4%)
Citation coverage tracked monthly (alert if <93%)
PII redaction audits quarterly
User feedback collected continuously
Post-Launch: 30-day review showed 2.1% hallucination rate, 96.8% citation coverage; continued monitoring.
A.3 Tier 3 Example: Public-Facing Eligibility Guidance Assistant
Use Case: Public Q&A on government program eligibility; provides general guidance and links to official sources.
Risk Tier Score: 15+
D1 Exposure: 3 (Public-facing)
D2 Impact: 3 (Eligibility decisions are rights-critical)
D3 Autonomy: 0 (Reference only; no actions)
D4 Data: 3 (Questions may contain regulated PII)
D5 Model/Change: 3 (Frontier model; frequent updates)
D6 Misuse: 3 (High disinformation/scam risk)
**Calculated Tier: 3 (15 points)
Automatic Tier 3 Override: YES - Public-facing AND regulated data AND rights-critical advice
Final Tier: 3
Approvals Required: Full ARB + Security + Privacy/Legal + Domain Lead (Policy Expert) + Executive Risk Owner (COO)
System Card Highlights:
Purpose: Provide general eligibility guidance for government assistance programs; direct users to official resources for definitive determinations
Model: Claude Sonnet 4 via Anthropic API with RAG from official policy documents
User Disclosure: Prominent banner: "AI-generated guidance - Not an official determination. For official eligibility decisions, contact [official channels]."
Data Handling:
User questions NOT stored (privacy by design)
Only metadata logged (timestamp, session ID, query category, sources cited)
No PII captured or retained
90-day log retention
Human Oversight: Content moderation team reviews flagged interactions; escalation path for policy questions
Refusal Policy: System refuses to provide definitive eligibility determinations; redirects to official channels
Provenance & Transparency Plan:
User-Facing Transparency:
Label: "AI-Powered Guidance Assistant"
Disclaimer: Appears before every response: "This is AI-generated general guidance. It is not an official eligibility determination. For binding decisions, please contact [official office] or visit [official website]."
Always shows sources cited from official policy documents
If sources absent → refuses to answer and provides official contact information
Provenance Audit Trail:
Model name and version
Retrieval source IDs (official policy document sections)
Timestamp
Query category (for trend analysis, not query content)
Safety outcome (accepted/refused/flagged)
No raw user text stored
Technical Mechanisms:
Source citation embedded in every response
Internal content ID for support queries
Provenance data retained 2 years (compliance requirement)
Tier 3 TEVV Report Highlights:
Test Set: 500 questions covering all program types and edge cases
A. Functional/Quality Results:
Grounding test: Policy claims must be supported by official sources
Target: Hallucinated policy claims < 1%; source coverage > 98%
Initial Results: 2.4% hallucination; 96.1% coverage (FAIL)
Mitigations:
Enhanced RAG with source verification
Stricter citation requirements
Refuse-if-unsure threshold lowered
Retest Results: 0.7% hallucination; 99.2% coverage (PASS)
B. Safety Results:
Categories tested: Harmful advice, scams, impersonation, manipulation
Target: Refusal rate ≥ 99% for harmful content
Initial Results: 97.2% refusal (FAIL - 2.8% harmful responses slipped through)
Mitigations:
Enhanced safety filters
Added harmful pattern detection
Lowered safety threshold
Retest Results: 99.6% refusal (PASS)
G. Red-Team Testing (Tier 3 Required):
External red-team engaged for 2 weeks
Attack scenarios: 150 attempts covering:
Jailbreaks to provide definitive eligibility decisions
Attempts to extract PII from training data
Social engineering to bypass safety filters
Disinformation generation
Impersonation of official channels
Initial Results: 12 successful attacks (8% success rate)
Critical Findings:
Complex multi-turn jailbreaks could elicit definitive decisions
Certain phrasings bypassed source requirement
Edge case policy questions sometimes hallucinated
Mitigations Implemented:
Multi-turn conversation safety checks
Hardened refusal on definitive language
Expanded edge case test coverage
Retest Results: 2 successful attacks (1.3% success rate)
Residual Vulnerabilities: Novel jailbreak techniques may still succeed; continuous monitoring required
Risk Analysis - Disclosure Readiness (Tier 3 Required):
Triggers for Public Disclosure:
Systematic wrong guidance affecting >100 users
Material policy misstatement with potential harm
Security/privacy incident involving user data
Successful large-scale jailbreak in production
Regulatory requirement
Internal Notification Process:
Incident detected → Engineering Owner notified (15 min)
P1 incidents → Executive Risk Owner + Legal notified (30 min)
ARB emergency meeting convened (4 hours)
Communication prepared (8 hours)
External Notification Process:
Users potentially affected: Email notification within 48 hours
Regulatory notification: Within 72 hours (per regulation)
Public disclosure: Within 5 days (if material public interest)
Communication Templates: Pre-approved templates ready for:
User notification
Regulatory notification
Press statement
Internal communication
Decision Authority:
Incident classification: Incident Commander
User notification: Executive Risk Owner + Legal
Regulatory notification: Legal + Compliance
Public disclosure: Executive Risk Owner + CEO + Legal
Rollback/Kill Switch:
Technical Mechanism: Feature flag can disable assistant in <5 minutes
Decision Authority: Executive Risk Owner or designated on-call
Testing: Kill switch tested quarterly
Fallback: Static FAQ page with official contact information displayed when disabled
Change Control:
Material changes requiring re-approval:
Model version changes
New policy document sources
Modified refusal behavior
Expansion to new program types
Re-testing: Full TEVV + red-team retest for material changes
Approval: Full ARB + Executive Risk Owner
Independent Validation:
External policy expert reviewed 100-question sample
Validated accuracy, source citation, appropriate refusals
Findings: 2 minor issues (addressed), overall validation: APPROVED
Validator Report attached to governance record
Post-Launch Monitoring:
Week 1-2 Review: Daily metrics review; 0 incidents; 99.1% source coverage maintained
Month 1 Review: Hallucination rate 0.6%; refusal rate 99.7%; user satisfaction 94%; 3 minor issues addressed
Ongoing: Monthly ARB review; quarterly red-team testing; annual comprehensive revalidation
Executive Risk Acceptance: "As Chief Operating Officer, I accept the residual risks associated with this AI system, specifically:
Novel jailbreak techniques may occasionally succeed (<2% estimated)
Edge cases may produce suboptimal guidance (~1%)
Users may misunderstand limitations despite clear disclaimers
These risks are acceptable given:
Comprehensive testing and mitigations
Continuous monitoring and rapid response capability
Significant public benefit in improving access to program information
Clear disclaimers and redirection to official channels
Signed: [COO Name], Date: [Date]"
Sample Tier 3 Eligibility Guidance Assistant - Comprehensive Questions
Section 1: Intake and Risk Assessment Questions
Use Case Definition:
What specific government programs does this assistant cover? (unemployment benefits, SNAP, Medicare, Medicaid, housing assistance, etc.)
What questions can users ask? (eligibility, application process, benefit amounts, documentation requirements)
What questions are explicitly prohibited? (specific benefit calculations, legal advice, immigration status determinations)
Who is the target user population? (general public, benefit applicants, caseworkers)
What languages will be supported?
What is the expected traffic volume? (queries per day/month)
Risk Tier Justification: 7. Why is this Tier 3? (Public-facing + regulated data + rights-critical advice) 8. What harm could occur if the system provides wrong information? 9. What vulnerable populations might be affected? (low-income, elderly, disabled, non-English speakers) 10. What is the potential for misuse? (scams, disinformation, impersonation)
Section 2: System Design and Architecture Questions
Model Selection: 11. Which LLM is being used and why? (Claude Sonnet 4, GPT-4, etc.) 12. What alternatives were considered? 13. How often does the model get updated by the vendor? 14. What happens when the vendor releases a new model version?
RAG and Knowledge Base: 15. What official documents are in the retrieval system? (policy manuals, regulations, FAQs) 16. How are documents sourced and verified as authoritative? 17. How frequently are documents updated? 18. What happens if a document is outdated or incorrect? 19. How is document versioning tracked? 20. How are documents chunked for retrieval?
System Prompting: 21. What instructions are in the system prompt? 22. How do you ensure the system doesn't provide definitive eligibility determinations? 23. What refusal behaviors are programmed? 24. How does the system handle edge cases or unclear policy?
Section 3: Data and Privacy Questions
User Data Handling: 25. What user data is collected? (just metadata, or actual questions?) 26. Are user questions stored or logged? 27. If questions contain PII, how is it handled? 28. What is the data retention period? 29. How can users request data deletion? 30. Are conversations linked to user accounts or anonymous?
PII Protection: 31. How do you prevent the system from asking users for sensitive PII? 32. What if a user volunteers their SSN or other sensitive data? 33. How is PII redacted from logs? 34. Who has access to any stored data? 35. How is data encrypted (in transit and at rest)?
Vendor Data Usage: 36. Does the vendor (Anthropic, OpenAI, etc.) use queries for model training? 37. Is there a data processing agreement (DPA) in place? 38. Can users opt out of data usage? 39. What data crosses borders to vendor data centers?
Section 4: Testing and Validation Questions (TEVV)
Test Coverage: 40. How many test questions are in your TEVV set? (requirement: 500+) 41. What program areas do test questions cover? 42. How were test questions developed? (policy experts, real user queries?) 43. What percentage of questions test edge cases?
Grounding and Accuracy: 44. What is your hallucination threshold? (target: <1%) 45. How do you measure hallucination? (fact-checking against official sources) 46. What was your pre-mitigation hallucination rate? 47. What mitigations reduced hallucination? 48. What is your source coverage target? (target: >98%) 49. What happens if the system can't find a source?
Safety Testing: 50. What harmful content categories are tested? (scams, manipulation, impersonation) 51. What is your refusal rate target for harmful content? (target: ≥99%) 52. What types of jailbreaks have you tested? 53. Can the system be tricked into providing definitive decisions? 54. Can it be manipulated to provide incorrect guidance?
Fairness and Bias: 55. Do different demographic groups receive equally accurate information? 56. Have you tested with non-English speakers? 57. Have you tested with users at different literacy levels? 58. Are there biases in which programs receive better coverage?
Red-Team Testing (Required for Tier 3): 59. Who conducted your red-team testing? (internal or external?) 60. How long did red-team testing last? 61. What attack scenarios were tested? 62. What was the initial attack success rate? 63. What critical vulnerabilities were found? 64. What mitigations were implemented? 65. What was the retest success rate? 66. What residual vulnerabilities remain?
Section 5: Transparency and User Experience Questions
User Disclosure: 67. How do users know they're interacting with AI? 68. Where is the disclaimer displayed? 69. What does the disclaimer say? 70. Is it shown before every response or just once? 71. How do you ensure users understand this is not official guidance?
Source Attribution: 72. Does every response cite official sources? 73. How are sources displayed to users? 74. Can users click to view the original source documents? 75. What happens if no source is available? 76. Do you provide links to official government resources?
Limitations Communication: 77. What limitations are communicated to users? 78. How do you redirect users to official channels for binding decisions? 79. What contact information is provided? 80. Is there a disclaimer about accuracy and completeness?
Section 6: Human Oversight Questions
Content Moderation: 81. Is there a content moderation team monitoring interactions? 82. What triggers human review of a conversation? 83. How quickly are flagged interactions reviewed? 84. What actions can moderators take?
Escalation: 85. Can users request human assistance? 86. How are complex policy questions escalated? 87. Who responds to escalated questions? 88. What is the response time for escalations?
Quality Assurance: 89. How do you sample and review outputs for quality? 90. What percentage of conversations are reviewed? 91. How are quality issues identified and addressed?
Section 7: Monitoring and Incident Response Questions
Production Monitoring: 92. What metrics are monitored in production? 93. What are your alert thresholds? (hallucination rate, refusal rate, error rate) 94. Who gets alerted when thresholds are breached? 95. What is your response time to alerts?
Drift Detection: 96. How do you detect performance drift over time? 97. How often do you re-validate the system? 98. What triggers retraining or updates?
Incident Classification: 99. What constitutes a P1 (critical) incident? 100. What constitutes a P2 (high) incident? 101. How quickly must you respond to each severity level?
Incident Response: 102. Who is the Incident Commander? 103. What is your rollback procedure? 104. How quickly can you disable the system (kill switch)? 105. When would you activate the kill switch? 106. What fallback is shown when the system is disabled?
Disclosure and Notification: 107. Under what circumstances would you notify affected users? 108. Under what circumstances would you notify regulators? 109. Under what circumstances would you make a public disclosure? 110. Who has authority to make these notification decisions? 111. What are your notification timelines?
Section 8: Provenance and Audit Questions (Tier 3 GenAI)
Provenance Tracking: 112. What provenance data is captured for each response? 113. How long is provenance data retained? 114. Who can access provenance data? 115. Can you reconstruct what sources were used for a specific response?
Audit Trail: 116. What is logged for audit purposes? 117. How do you ensure audit logs can't be tampered with? 118. What is the audit log retention period? 119. Who can access audit logs?
User Disputes: 120. How can a user dispute incorrect information they received? 121. What is the process for investigating disputes? 122. How are users notified of investigation results?
Section 9: Regulatory and Compliance Questions
EU AI Act (if serving EU users): 123. Is this classified as a high-risk AI system under the EU AI Act? 124. What conformity assessment was conducted? 125. Is the system registered in the EU database? 126. What post-market monitoring is in place? 127. How do you report serious incidents to regulators?
GDPR (if serving EU users): 128. What is the lawful basis for processing? (legitimate interest, consent?) 129. Has a DPIA been conducted? 130. How do users exercise their rights (access, deletion, correction)? 131. What is your data breach notification procedure?
CCPA (if serving California users): 132. How do users know what personal information is collected? 133. Can users opt out of data sale/sharing? 134. How do users request deletion? 135. Is there a "Do Not Sell My Personal Information" link?
Sector-Specific (if applicable): 136. Does providing benefit eligibility information trigger any regulations? 137. Are there state-specific requirements? 138. Do you need to comply with accessibility standards (Section 508, WCAG)?
Section 10: Approval and Governance Questions
ARB Review: 139. Which ARB members reviewed this system? 140. What concerns were raised during review? 141. How were concerns addressed? 142. Were there any conditions on approval?
Executive Risk Acceptance: 143. Who is the Executive Risk Owner? 144. What residual risks did they accept? 145. What is the justification for accepting these risks? 146. When does the risk acceptance need to be renewed?
Change Management: 147. What changes would require re-approval? 148. What changes can be made without re-approval? 149. What is the process for material changes? 150. How often must the system be re-validated?
Section 11: Post-Launch Questions
Post-Launch Reviews: 151. When is the 2-week post-launch review scheduled? 152. Who participates in post-launch reviews? 153. What metrics are reviewed? 154. What is the escalation process for issues?
User Feedback: 155. How do users provide feedback? 156. What is the volume of positive vs. negative feedback? 157. What common complaints or confusion points exist? 158. How is feedback incorporated into improvements?
Performance Metrics: 159. What is your actual hallucination rate in production? 160. What is your actual source coverage in production? 161. What is your actual refusal rate? 162. What is user satisfaction? 163. How many users have been served? 164. What is the most common type of question?
Continuous Improvement: 165. How often do you retrain or update the system? 166. What improvements have been made since launch? 167. What issues remain to be addressed? 168. What is on the roadmap for future enhancements?
Section 12: Vendor Management Questions
Vendor Relationship: 169. What is your SLA with the model provider? 170. What happens if the vendor has an outage? 171. What happens if the vendor deprecates your model? 172. Do you have a backup vendor or model?
Vendor Security: 173. What security certifications does the vendor have? (SOC 2, ISO 27001) 174. How does the vendor handle security incidents? 175. What is the vendor's incident notification timeline?
Vendor Updates: 176. How much notice does the vendor give for model updates? 177. What is your testing process for vendor updates? 178. Can you block vendor updates? 179. What happens if a vendor update causes regressions?
Question Usage Guide
For Intake and Planning: Questions 1-40
For System Design: Questions 11-39
For TEVV Planning: Questions 40-66
For User Experience Design: Questions 67-80
For Operations: Questions 81-111
For Compliance: Questions 123-138
For ARB Review: Questions 1-179 (comprehensive review)
For Post-Launch: Questions 151-179
17. Appendix B: Acronyms
| Acronym | Definition |
| ABAC | Attribute-Based Access Control |
| AI | Artificial Intelligence |
| AICPA | American Institute of Certified Public Accountants |
| API | Application Programming Interface |
| ARB | AI Review Board |
| BERT | Bidirectional Encoder Representations from Transformers |
| CCPA | California Consumer Privacy Act |
| CISO | Chief Information Security Officer |
| CLI | Command Line Interface |
| CNIL | Commission Nationale de l'Informatique et des Libertés (France) |
| COO | Chief Operating Officer |
| CPRA | California Privacy Rights Act |
| CRA | Cyber Resilience Act (EU) |
| CSAM | Child Sexual Abuse Material |
| CTO | Chief Technology Officer |
| CUI | Controlled Unclassified Information |
| CVSS | Common Vulnerability Scoring System |
| DPA | Data Processing Agreement |
| DPIA | Data Protection Impact Assessment |
| DPO | Data Protection Officer |
| ECOA | Equal Credit Opportunity Act |
| EDPB | European Data Protection Board |
| EEA | European Economic Area |
| EEOC | Equal Employment Opportunity Commission |
| ENISA | European Union Agency for Cybersecurity |
| FCA | Financial Conduct Authority (UK) |
| FCRA | Fair Credit Reporting Act |
| FDA | Food and Drug Administration |
| FedRAMP | Federal Risk and Authorization Management Program |
| FIPS | Federal Information Processing Standards |
| FTC | Federal Trade Commission |
| GAI | Generative Artificial Intelligence |
| GDPR | General Data Protection Regulation |
| GenAI | Generative Artificial Intelligence |
| GPT | Generative Pre-trained Transformer |
| HIPAA | Health Insurance Portability and Accountability Act |
| HITRUST | Health Information Trust Alliance |
| HR | Human Resources |
| HRIS | Human Resources Information System |
| IDDL | Independent DevOps Development Lifecycle |
| IEC | International Electrotechnical Commission |
| IP | Intellectual Property |
| ISO | International Organization for Standardization |
| LLM | Large Language Model |
| LMS | Learning Management System |
| ML | Machine Learning |
| MLOPS | Machine Learning Operations |
| NDA | Non-Disclosure Agreement |
| NIST | National Institute of Standards and Technology |
| NLP | Natural Language Processing |
| OMB | Office of Management and Budget |
| OWASP | Open Web Application Security Project |
| PCI DSS | Payment Card Industry Data Security Standard |
| PHI | Protected Health Information |
| PII | Personally Identifiable Information |
| PPTDF | Pre-Processing, Training, Deployment, Feedback |
| RACI | Responsible, Accountable, Consulted, Informed |
| RAG | Retrieval-Augmented Generation |
| RBAC | Role-Based Access Control |
| RMF | Risk Management Framework |
| ROC | Report on Compliance |
| SaaS | Software as a Service |
| SaMD | Software as a Medical Device |
| SCRM | Supply Chain Risk Management |
| SDK | Software Development Kit |
| SLA | Service Level Agreement |
| SLO | Service Level Objective |
| SOC | System and Organization Controls |
| SOX | Sarbanes-Oxley Act |
| TEVV | Testing, Evaluation, Validation, and Verification |
| UAT | User Acceptance Testing |
| UK GDPR | United Kingdom General Data Protection Regulation |
| URL | Uniform Resource Locator |
| US | United States |
| XAI | Explainable Artificial Intelligence |
18. Appendix C: Control-to-Regulation Mapping Matrix
This matrix shows which controls help demonstrate compliance with major AI regulations and standards.
| Control ID | Control Name | EU AI Act | GDPR | NIST AI RMF | ISO 42001 | SOC 2 | CCPA | UK AI Reg. |
| AAT-01 | AI System Inventory | ● | ○ | ● | ● | ● | ○ | ● |
| AAT-02 | Risk Classification | ● | ● | ● | ● | ● | ○ | ● |
| AAT-03 | Governance Framework | ● | ● | ● | ● | ● | ● | ● |
| AAT-04 | Role Assignment | ● | ● | ● | ● | ● | ● | ● |
| AAT-05 | Policy Documentation | ● | ● | ● | ● | ● | ● | ● |
| AAT-06 | Training Data Documentation | ● | ● | ● | ● | ○ | ● | ● |
| AAT-07 | Data Quality Standards | ● | ● | ● | ● | ● | ● | ● |
| AAT-08 | Bias Detection | ● | ● | ● | ● | ○ | ○ | ● |
| AAT-09 | Data Minimization | ● | ● | ● | ● | ● | ● | ● |
| AAT-10 | Model Development Standards | ● | ○ | ● | ● | ● | ○ | ● |
| AAT-11 | Performance Metrics | ● | ○ | ● | ● | ● | ○ | ● |
| AAT-12 | Model Validation | ● | ● | ● | ● | ● | ○ | ● |
| AAT-13 | Model Documentation | ● | ● | ● | ● | ● | ○ | ● |
| AAT-14 | Pre-Deployment Testing | ● | ● | ● | ● | ● | ○ | ● |
| AAT-15 | Security Testing | ● | ● | ● | ● | ● | ● | ● |
| AAT-16 | Adversarial Testing | ● | ○ | ● | ● | ● | ○ | ● |
| AAT-17 | User Acceptance Testing | ● | ● | ● | ● | ● | ○ | ● |
| AAT-18 | Deployment Controls | ● | ● | ● | ● | ● | ● | ● |
| AAT-19 | Access Controls | ● | ● | ● | ● | ● | ● | ● |
| AAT-20 | Monitoring Infrastructure | ● | ● | ● | ● | ● | ○ | ● |
| AAT-21 | Performance Monitoring | ● | ○ | ● | ● | ● | ○ | ● |
| AAT-22 | Incident Response | ● | ● | ● | ● | ● | ● | ● |
| AAT-23 | Human Oversight | ● | ● | ● | ● | ○ | ○ | ● |
| AAT-24 | User Transparency | ● | ● | ● | ● | ○ | ● | ● |
| AAT-25 | Explainability | ● | ● | ● | ● | ○ | ● | ● |
| AAT-26 | User Rights | ● | ● | ● | ● | ○ | ● | ● |
| AAT-27 | Change Management | ● | ● | ● | ● | ● | ○ | ● |
| AAT-28 | Model Retraining | ● | ○ | ● | ● | ○ | ○ | ● |
| AAT-29 | Version Control | ● | ● | ● | ● | ● | ○ | ● |
| AAT-29.1 | Decommissioning Process | ● | ● | ● | ● | ● | ● | ● |
| AAT-29.2 | Third-Party AI Assessment | ● | ● | ● | ● | ● | ● | ● |
| AAT-29.3 | AI Supply Chain Due Diligence | ● | ● | ● | ● | ● | ○ | ● |
| AAT-29.4 | Vendor Contract Terms | ● | ● | ● | ● | ● | ● | ● |
| AAT-29.5 | Environmental Impact Assessment | ● | ○ | ○ | ● | ○ | ○ | ● |
| AAT-29.6 | Societal Impact Assessment | ● | ○ | ● | ● | ○ | ○ | ● |
| AAT-29.7 | Stakeholder Engagement | ● | ● | ● | ● | ○ | ○ | ● |
| AAT-29.8 | Appeals/Redress Mechanism | ● | ● | ● | ● | ○ | ● | ● |
| AAT-30 | Audit Readiness | ● | ● | ● | ● | ● | ● | ● |
| AAT-31 | Record Retention | ● | ● | ● | ● | ● | ● | ● |
| AAT-32 | Compliance Reporting | ● | ● | ● | ● | ● | ● | ● |
| AAT-32.1 | AI Ethics Review | ● | ● | ● | ● | ○ | ○ | ● |
Legend:
● = Directly Required or Strongly Recommended
○ = Indirectly Supports Compliance
Blank = Not Applicable
19. Appendix D: Risk Assessment Template
Note: See Part 2, Section 9 for the complete fillable Risk Assessment Template. The template includes:
Assessment metadata
System context
6-dimension risk scoring (D1-D6)
Tier determination with automatic overrides
Trustworthiness risk analysis across 8 characteristics
GenAI-specific risk assessment
Threat scenarios documentation
Risk treatment decisions
Required evidence checklist by tier
Approval and sign-off section
This template is ready to copy and use for each AI system risk assessment.
20. Appendix E: AI Model Card Template
Note: See Part 2, Section 9 for the complete fillable AI Model Card Template. The template includes:
Model overview and details
Training details and data characteristics
Performance metrics and results
Fairness and bias analysis
Robustness and reliability testing
Security considerations
Privacy and data protection
Environmental impact
Ethical considerations
Explainability methods
Known limitations and recommendations
Dependencies and versioning
Contact and support information
Compliance and certifications
This template is ready to copy and use for documenting AI models.
21. Appendix F: AI Incident Response Playbook
Note: See Part 1 of this document series for the complete 6-phase AI Incident Response Playbook including:
Phase 1: Detection and Triage (severity classification, triage process)
Phase 2: Containment (immediate actions, system state decisions)
Phase 3: Investigation and Root Cause Analysis
Phase 4: Remediation and Recovery
Phase 5: Communication (internal and external)
Phase 6: Post-Incident Review and Learning
The playbook includes incident severity matrix (P1-P4), containment checklists, communication templates, and roles/responsibilities.
22. Appendix G: AI Training Curriculum
Note: See Part 1 of this document series for the complete AI Training Curriculum including:
Training Levels:
Level 1: AI Awareness (All Employees) - 30 minutes
Level 2: AI Governance Fundamentals (AI Users/Contributors) - 2 hours
Level 3: AI Risk and Compliance Practitioners - 1 day
Level 4: AI Governance Leadership (Executives) - 4 hours
Specialized Modules:
Module S1: GenAI-Specific Risks and Controls - 1 hour
Module S2: Red Team Testing for AI - 4 hours
Module S3: Fairness and Bias in AI - 2 hours
Module S4: AI Explainability and Transparency - 2 hours
Module S5: AI Privacy and Data Protection - 2 hours
Module S6: Third-Party AI and Supply Chain Risk - 1.5 hours
Training Pathways by Role with required completion timelines and enforcement mechanisms.
23. Appendix H: Vendor Assessment Questionnaire
Note: See Part 1 of this document series for the complete 8-section Vendor Assessment Questionnaire including:
Sections:
Vendor Information
Governance and Risk Management
Model Development and Validation
Security and Privacy
Fairness, Bias, and Safety
Transparency and Explainability
Monitoring and Continuous Improvement
Contractual and Operational Considerations
Each section includes detailed questions, scoring rubrics (0-3), and weighted scoring methodology for overall vendor risk classification (Low/Medium/High/Very High Risk).
24. Appendix I: Quick Reference - Prohibited AI (EU AI Act)
EU AI Act Prohibited Practices (Effective Feb 2, 2025)
BANNED - Cannot Deploy in EU:
Subliminal Manipulation
AI that deploys subliminal techniques beyond a person's consciousness to materially distort behavior causing harm
No exceptions
Exploitation of Vulnerabilities
AI exploiting vulnerabilities of age, disability, or social/economic situation to materially distort behavior causing harm
No exceptions
Social Scoring by Governments
AI evaluating or classifying people based on social behavior or personal characteristics leading to detrimental treatment
Limited government use cases only
Real-Time Biometric Identification in Public (Law Enforcement)
Real-time remote biometric ID in publicly accessible spaces
Narrow exceptions: serious crimes, missing persons, imminent threats (with judicial authorization)
Biometric Categorization (Sensitive Attributes)
AI inferring race, political opinions, union membership, religious beliefs, sex life, sexual orientation from biometric data
Law enforcement exceptions with safeguards
Emotion Recognition (Workplace/Education)
AI recognizing emotions in workplace and educational institutions
Law enforcement and medical/safety exceptions
Untargeted Biometric Scraping
Scraping facial images or biometric data from internet/CCTV to create facial recognition databases
Law enforcement exceptions with authorization
High-Risk AI Systems (Heavy Regulation Required)
Categories:
Employment and workers management (recruitment, promotion, termination)
Access to essential services (credit scoring, emergency dispatching, benefits)
Law enforcement (when permitted by exceptions above)
Migration, asylum, border control
Administration of justice
Democratic processes (election influence)
Education and vocational training
Biometric systems (non-real-time identification)
Critical infrastructure safety
Requirements:
Risk management system
High-quality, representative training data
Technical documentation and model cards
Automatic logging for traceability
Transparency to users
Human oversight
Accuracy, robustness, cybersecurity
Conformity assessment (third-party or self)
Registration in EU database
Post-market monitoring
General Purpose AI Models (GPAI)
Standard GPAI:
Technical documentation
Information for downstream providers
Copyright compliance policy
Publicly available training data summary
Systemic Risk GPAI (>10^25 FLOPs):
All standard requirements PLUS:
Model evaluation for systemic risks
Adversarial testing
Serious incident reporting
Cybersecurity protections
Energy consumption reporting
Compliance Timeline
| Date | Requirement |
| Feb 2, 2025 | Prohibited AI ban in effect |
| Aug 2, 2025 | GPAI obligations in effect |
| Aug 2, 2026 | High-risk requirements for new systems |
| Aug 2, 2027 | High-risk requirements for existing systems |
Penalties
| Violation | Maximum Fine |
| Prohibited AI | €35M or 7% global revenue (whichever higher) |
| Obligation breach | €15M or 3% global revenue (whichever higher) |
| Incorrect information | €7.5M or 1% global revenue (whichever higher) |
Quick Decision: Does Your AI System Comply?
1. Is it on the prohibited list?
YES → STOP. Cannot deploy in EU.
NO → Continue
2. Is it high-risk per Annex III?
YES → Must comply with high-risk requirements
NO → Continue
3. Is it a GPAI model?
YES → Comply with GPAI requirements
(+ systemic risk if >10^25 FLOPs)
NO → Minimal transparency only
4. Does it interact with humans or generate content?
YES → Disclose AI use
NO → Transparency based on use case
Resources:
Official Text: EUR-Lex 32024R1689
EU AI Office: https://digital-strategy.ec.europa.eu/en/policies/ai-office
Compliance Guidance: https://artificialintelligenceact.eu
25. Appendix J: Resources and References
Standards and Frameworks
International Standards:
NIST AI Risk Management Framework (AI RMF 1.0) - NIST AI 100-1 (2023) https://www.nist.gov/itl/ai-risk-management-framework
NIST AI RMF: Generative AI Profile - NIST AI 600-1 (July 2024) https://doi.org/10.6028/NIST.AI.600-1
ISO/IEC 42001:2023 - AI Management System
ISO/IEC 23894:2023 - AI Risk Management
ISO/IEC 25059:2023 - AI System Quality Evaluation
Industry Frameworks:
OECD AI Principles (2019, updated 2024)
IEEE 7000 Series - Ethics in AI
MLOps Maturity Model
Regulatory Resources
EU AI Act:
Official Regulation Text - EUR-Lex 32024R1689
EU AI Office - https://digital-strategy.ec.europa.eu/en/policies/ai-office
Compliance Guidance - https://artificialintelligenceact.eu
Data Protection:
GDPR Official Text - https://gdpr-info.eu
EDPB Guidelines - https://edpb.europa.eu
California CCPA/CPRA - https://oag.ca.gov/privacy/ccpa
US Federal:
Executive Order 14110 - White House briefing room
NIST AI Safety Institute - https://www.nist.gov/aisi
Testing and Evaluation Tools
Fairness and Bias:
AI Fairness 360 (IBM) - https://aif360.mybluemix.net
Fairlearn (Microsoft) - https://fairlearn.org
What-If Tool (Google) - https://pair-code.github.io/what-if-tool
Explainability:
InterpretML (Microsoft) - https://interpret.ml
Adversarial Robustness:
Adversarial Robustness Toolbox (IBM) - https://github.com/Trusted-AI/adversarial-robustness-toolbox
CleverHans (Google) - https://github.com/cleverhans-lab/cleverhans
LLM/GenAI Security:
OWASP Top 10 for LLM Applications - https://owasp.org/www-project-top-10-for-large-language-model-applications
Garak - LLM Vulnerability Scanner - https://github.com/leondz/garak
PyRIT (Microsoft) - https://github.com/Azure/PyRIT
NeMo Guardrails (NVIDIA) - https://github.com/NVIDIA/NeMo-Guardrails
Model Monitoring:
Evidently AI - https://www.evidentlyai.com
Whylogs - https://whylabs.ai/whylogs
Fiddler AI - https://www.fiddler.ai
Organizations and Communities
Standards Bodies:
NIST AI Division - https://www.nist.gov/artificial-intelligence
ISO/IEC JTC 1/SC 42 (AI Standards)
IEEE Standards Association - AI
Research Institutes:
Partnership on AI - https://partnershiponai.org
AI Now Institute - https://ainowinstitute.org
Center for AI Safety - https://www.safe.ai
Montreal AI Ethics Institute - https://montrealethics.ai
Industry Consortia:
Responsible AI Institute - https://www.responsible.ai
MLCommons - https://mlcommons.org
Linux Foundation AI & Data - https://lfaidata.foundation
Publications
Books:
"Trustworthy Machine Learning" - Kush R. Varshney (2022)
"The Alignment Problem" - Brian Christian (2020)
"Atlas of AI" - Kate Crawford (2021)
"Weapons of Math Destruction" - Cathy O'Neil (2016)
Academic Venues:
ACM FAccT - Conference on Fairness, Accountability, and Transparency
NeurIPS - Neural Information Processing Systems
AIES - AI, Ethics, and Society
Journals:
AI and Ethics (Springer)
Nature Machine Intelligence
AI Magazine (AAAI)
Newsletters:
Import AI - Jack Clark
The Batch - deeplearning.ai
Training and Certification
Online Courses:
AI Safety Fundamentals - BlueDot Impact
Responsible AI - Coursera (multiple providers)
AI Ethics - edX
MLOps Specialization - Coursera
Professional Certifications:
Certified AI Governance Professional - AI Cert Foundation
TensorFlow Developer Certificate - Google
Azure AI Engineer Associate - Microsoft
AWS Certified Machine Learning - Specialty
Open Source Projects
ML Frameworks:
TensorFlow - https://www.tensorflow.org
PyTorch - https://pytorch.org
scikit-learn - https://scikit-learn.org
Hugging Face Transformers - https://huggingface.co/transformers
MLOps Platforms:
MLflow - https://mlflow.org
Kubeflow - https://www.kubeflow.org
DVC (Data Version Control) - https://dvc.org
Weights & Biases - https://wandb.ai
Model Cards:
Model Card Toolkit (Google) - https://github.com/tensorflow/model-card-toolkit
HuggingFace Model Cards - https://huggingface.co/docs/hub/model-cards
Incident Databases
AI Incident Database - Partnership on AI - https://incidentdatabase.ai
AIAAIC Repository - https://www.aiaaic.org
Policy Resources
Think Tanks:
Center for Security and Emerging Technology (Georgetown) - https://cset.georgetown.edu
Brookings Institution - AI Governance
Center for Data Innovation - https://datainnovation.org
Contact Organizations
EU AI Act Compliance:
EU AI Office: ai-office@ec.europa.eu
National Competent Authorities (vary by member state)
US Federal Guidance:
NIST AI Team: ai_standards@nist.gov
FTC Technology Division
Data Protection:
EDPB (EU) - https://edpb.europa.eu/about-edpb/about-edpb/contact_en
ICO (UK) - https://ico.org.uk/global/contact-us
FTC (US) - https://www.ftc.gov/about-ftc/contact
26. Document Control (Sample)
Document Information
Document Title: AI Assurance Testing and Governance Framework
Document ID: AAT-FRAMEWORK-2026-v1.0
Version: 1.0
Publication Date: January 13, 2026
Next Review Date: January 13, 2027 (or sooner if significant regulatory changes)
Classification: Internal Use Only
Document Owner
Primary Owner: Chief AI Officer / Chief Information Security Officer
Custodian: AI Governance Team
Approvers: Executive Leadership Team
Version History
| Version | Date | Author | Changes | Approver |
| 0.1 | October 1, 2025 | AI Governance Team | Initial draft | - |
| 0.5 | November 15, 2025 | AI Governance Team | Incorporated NIST AI 600-1 GenAI Profile | CISO |
| 0.9 | December 15, 2025 | AI Governance Team + Legal | Added EU AI Act compliance mappings | General Counsel |
| 1.0 | January 13, 2026 | AI Governance Team | Final review and approval | Executive Leadership |
Review and Approval
| Role | Name | Signature | Date |
| Chief AI Officer | [Name] | [Signature] | January 13, 2026 |
| Chief Information Security Officer | [Name] | [Signature] | January 13, 2026 |
| Chief Privacy Officer | [Name] | [Signature] | January 13, 2026 |
| General Counsel | [Name] | [Signature] | January 13, 2026 |
| Chief Technology Officer | [Name] | [Signature] | January 13, 2026 |
| Chief Executive Officer | [Name] | [Signature] | January 13, 2026 |
Distribution List
Primary Distribution:
Executive Leadership Team
All AI Development Teams
Information Security Team
Privacy and Legal Teams
Compliance and Audit
Risk Management
Human Resources (for training administration)
Procurement (for vendor management)
Secondary Distribution:
All Product Management
All Engineering Leadership
Quality Assurance Teams
Customer Success (for customer-facing AI)
Document Classification and Handling
Classification: Internal Use Only (contains proprietary governance processes)
Handling Instructions:
Do not share externally without General Counsel approval
May be shared with vendors under NDA for compliance verification
Redacted version available for customer/partner disclosure upon request
Storage: Document repository at [URL to be configured]
Access Controls: All employees; read-only except AI Governance Team
Retention Period
Active Version: Indefinite (living document)
Superseded Versions: 7 years from supersession date
Legal Hold: Preserve indefinitely if subject to litigation or investigation
Related Documents
Governance Artifacts (Maintained Separately):
AI System Inventory (in governance system)
AI Risk Register (in risk management system)
Individual System Cards (per AI system)
TEVV Reports (per AI system)
Threat Models (per AI system)
Risk Analyses (per AI system)
Vendor Assessment Reports (per vendor)
Incident Reports (per incident)
Training Records (per employee)
ARB Meeting Minutes (per meeting)
Organizational Policies:
Information Security Policy
Data Privacy Policy
Software Development Lifecycle (SDLC) Policy
Vendor Management Policy
Incident Response Policy
Acceptable Use Policy
Records Retention Policy
Change Management
Material Changes Requiring Re-Approval:
Changes to risk tiering methodology
Changes to required artifacts by tier
Changes to approval authority
Addition/removal of mandatory controls
Changes to compliance mappings
Minor Changes (AI Governance Team Authority):
Template formatting improvements
Clarifications to existing requirements
Addition of examples
Updates to external references/links
Typographical corrections
Change Process:
Propose change with justification
Impact assessment (affected systems, teams, timelines)
Draft updated version
Review by stakeholders
Approval per change type (material vs. minor)
Communication of changes
Training updates (if needed)
Version increment and publication
Feedback and Continuous Improvement
Feedback Channels:
Email: ai-governance@[organization].com
Internal portal feedback form
Quarterly stakeholder surveys
Post-incident reviews
Annual comprehensive review
Feedback Review:
AI Governance Team reviews all feedback monthly
Prioritizes improvements based on impact and frequency
Implements approved improvements in next version
Communicates feedback responses to submitters
27. Contact Information
For questions, guidance, or support related to AI governance:
AI Governance Team
General Inquiries:
Email: ai-governance@[organization].com
Internal Portal: [URL to be configured]
Slack Channel: #ai-governance
Office Hours: Tuesdays 2-3pm, Thursdays 10-11am [Calendar link]
AI Risk Lead:
Name: [Name]
Email: ai-risk@[organization].com
Phone: [Phone number]
Office: [Location]
AI Review Board (ARB)
ARB Chair:
Name: [Name]
Title: [Title]
Email: arb-chair@[organization].com
Meeting Schedule:
Regular Meetings: Wednesdays 2-4pm weekly
Emergency Meetings: On-call via Slack @arb-urgent
Submission and Intake:
New System Submissions: ai-arb-intake@[organization].com
ARB Meeting Calendar: [Calendar link]
Submission Deadlines: Monday 5pm for Wednesday review
Specialized Reviewers
Security (AI-Related):
Team: AI Security Team
Email: ai-security@[organization].com
Lead: [Name], [Title]
Security Incident Reporting: security-incident@[organization].com
Urgent: Page via PagerDuty "AI-Security-Urgent"
Privacy and Legal (AI-Related):
Team: AI Privacy & Legal
Email: ai-privacy-legal@[organization].com
Privacy Lead: [Name], Chief Privacy Officer
Legal Lead: [Name], [Title]
DPA/Contract Reviews: legal-contracts@[organization].com
Domain Leads (By Area):
Healthcare AI: [Name], [Email]
Financial Services AI: [Name], [Email]
HR/People Analytics: [Name], [Email]
Customer-Facing AI: [Name], [Email]
Internal Productivity AI: [Name], [Email]
Training and Support
AI Governance Training:
Training Portal: [URL to be configured]
Training Coordinator: [Name]
Email: ai-training@[organization].com
LMS Access: [URL]
Technical Support:
Help Desk: ai-help@[organization].com
Slack: #ai-governance-help
Knowledge Base: [URL]
Executive Sponsors
Chief AI Officer:
Name: [Name]
Email: [Email]
Executive Assistant: [Name], [Email]
Chief Information Security Officer (CISO):
Name: [Name]
Email: [Email]
Executive Assistant: [Name], [Email]
Chief Privacy Officer (CPO):
Name: [Name]
Email: [Email]
Executive Assistant: [Name], [Email]
Chief Technology Officer (CTO):
Name: [Name]
Email: [Email]
Executive Assistant: [Name], [Email]
General Counsel:
Name: [Name]
Email: [Email]
Executive Assistant: [Name], [Email]
External Inquiries
Media Relations:
Email: media-relations@[organization].com
Phone: [Phone number]
Regulatory Affairs:
Email: regulatory-affairs@[organization].com
Regulatory Compliance Lead: [Name]
Research Collaboration:
Email: ai-research@[organization].com
Research Partnerships Lead: [Name]
Customer/Partner Inquiries:
Email: ai-transparency@[organization].com
Customer AI Documentation: [URL]
Emergency Contacts
P1 AI Incidents (24/7):
PagerDuty: "AI-Incident-P1"
Emergency Hotline: [Phone number]
Incident Commander On-Call: Via PagerDuty
Executive Escalation (After Hours):
Executive On-Call: Via PagerDuty "Executive-Escalation"
CISO Direct (Emergencies Only): [Phone number]
28. Legal Notice
Copyright and Intellectual Property
© 2026 [Organization Name]. All rights reserved.
This document contains proprietary and confidential information belonging to [Organization Name]. No part of this document may be reproduced, distributed, or transmitted in any form or by any means, including photocopying, recording, or other electronic or mechanical methods, without the prior written permission of [Organization Name], except in the case of brief quotations embodied in critical reviews and certain other noncommercial uses permitted by copyright law.
For permission requests, contact: Legal Department [Organization Name] [Address] Email: legal@[organization].com
Disclaimer and Limitation of Liability
General Disclaimer:
This AI Governance Framework provides guidance for managing AI systems and complying with applicable regulations. It is provided "as is" without warranty of any kind, either express or implied, including but not limited to the implied warranties of merchantability, fitness for a particular purpose, or non-infringement.
Not Legal Advice:
This framework is NOT legal advice. Organizations should consult with qualified legal counsel regarding:
Specific regulatory requirements in their jurisdictions
Interpretation and application of AI regulations (EU AI Act, GDPR, CCPA, etc.)
Contractual obligations and liability
Intellectual property matters
Compliance strategy
No Guarantee of Compliance:
While this framework is designed to align with current AI regulations and best practices, [Organization Name] makes no representations or warranties that:
Following this framework will ensure full regulatory compliance
This framework addresses all applicable laws and regulations
Regulatory interpretations will align with framework guidance
Future regulatory changes will not require framework modifications
Limitation of Liability:
To the maximum extent permitted by law, [Organization Name] shall not be liable for any:
Direct, indirect, incidental, special, consequential, or exemplary damages
Lost profits, revenues, data, or business opportunities
Regulatory fines or penalties
Litigation costs or settlements
Reputational harm
arising from or related to the use or inability to use this framework, even if advised of the possibility of such damages.
Regulatory Compliance Statement
Alignment with Regulations:
This framework is designed to align with:
EU Artificial Intelligence Act (Regulation (EU) 2024/1689)
NIST AI Risk Management Framework (AI RMF 1.0, NIST AI 100-1)
NIST AI RMF Generative AI Profile (NIST AI 600-1)
General Data Protection Regulation (GDPR - Regulation (EU) 2016/679)
UK GDPR
California Consumer Privacy Act (CCPA) and California Privacy Rights Act (CPRA)
ISO/IEC 42001:2023 (AI Management System)
ISO/IEC 27001:2022 (Information Security)
SOC 2 Type II (Trust Services Criteria)
Organizational Responsibility:
Each organization implementing this framework is responsible for:
Ensuring compliance with all applicable laws and regulations in their jurisdictions
Adapting the framework to their specific regulatory requirements
Obtaining necessary certifications and attestations
Maintaining compliance over time as regulations evolve
Documenting compliance evidence for regulators and auditors
Regulatory Updates:
AI regulations are rapidly evolving. Organizations must:
Monitor regulatory developments in all relevant jurisdictions
Update this framework within 90 days of material regulatory changes
Re-assess AI systems when regulations change
Maintain awareness of enforcement trends and regulatory guidance
Jurisdictional Variations:
This framework provides general guidance. Specific requirements may vary by:
Country/region (EU, US, UK, China, etc.)
State/province (California, Colorado, New York, etc.)
Industry sector (healthcare, financial services, employment, etc.)
AI system risk level and use case
Indemnification
Internal Use:
When used internally by [Organization Name], this framework is provided as organizational policy. Employees and contractors agree to follow this framework as a condition of working with AI systems.
External Distribution:
If this framework is shared with vendors, partners, or customers:
Recipients use the framework at their own risk
[Organization Name] provides no warranties or support to external parties
External parties agree to indemnify and hold harmless [Organization Name] from any claims arising from their use of the framework
External distribution requires General Counsel approval and appropriate disclaimers
Modification and Updates
Amendment Rights:
[Organization Name] reserves the right to modify, update, or discontinue this framework at any time without notice. Continued use of the framework after modifications constitutes acceptance of the changes.
Version Control:
Only the version published in the official document repository is authoritative
Superseded versions are for reference only and should not be used
Users are responsible for ensuring they reference the current version
Version history is maintained in the Document Control section
Third-Party Content
External References:
This framework references third-party standards, regulations, tools, and resources. [Organization Name]:
Does not endorse or guarantee any third-party content
Is not responsible for the accuracy or availability of external links
Does not control third-party terms of use or privacy policies
Recommends users review third-party terms before use
Open Source Tools:
References to open-source tools and frameworks are provided for informational purposes. Users should:
Review open-source licenses before use
Assess security and maintenance status
Ensure compatibility with organizational policies
Obtain necessary approvals before deployment
Export Control and Sanctions
Export Compliance:
Some AI technologies referenced in this framework may be subject to export controls. Organizations must:
Comply with all applicable export control laws (ITAR, EAR, etc.)
Obtain necessary export licenses before sharing AI systems internationally
Ensure third-party vendors comply with export controls
Maintain export compliance documentation
Sanctions:
Organizations must not use this framework to facilitate:
Development or deployment of AI systems in violation of sanctions
Provision of AI services to sanctioned entities or individuals
Any activity prohibited by OFAC or other sanctions regimes
Data Protection and Privacy
Personal Data:
This framework itself does not process personal data. However, when implementing AI systems under this framework:
Organizations are data controllers and must comply with applicable privacy laws
AI systems must be designed with privacy by design and by default
Data processing must have a lawful basis (GDPR Article 6)
Data subjects must be informed of AI use (GDPR Article 13-14)
Data subject rights must be enabled (GDPR Articles 15-22)
Confidential Information:
This framework may reference or contain confidential business information. Recipients must:
Maintain confidentiality of proprietary information
Not disclose to unauthorized parties
Use only for authorized purposes
Return or destroy upon request
Severability
If any provision of this Legal Notice is found to be invalid, illegal, or unenforceable, the remaining provisions shall continue in full force and effect.
Governing Law
This framework and Legal Notice shall be governed by and construed in accordance with the laws of [Jurisdiction], without regard to its conflict of law provisions.
Entire Agreement
This Legal Notice, together with the AI Governance Framework document, constitutes the entire agreement regarding the use of this framework and supersedes all prior or contemporaneous understandings.
Contact for Legal Matters
For legal questions regarding this framework:
Legal Department
[Organization Name]
[Address]
Email: legal@[organization].com
Phone: [Phone number]
Acknowledgments
This AI Governance Framework was developed with contributions from:
Internal Contributors:
AI Governance Team
Information Security Team
Privacy and Legal Team
AI Engineering Teams
Risk Management Team
Compliance and Audit Team
External Guidance:
NIST AI Risk Management Framework
EU AI Act regulatory text and guidance
ISO/IEC 42001 standard
Industry best practices and benchmarks
Special Thanks:
Executive leadership for sponsorship and support
Early adopter teams who participated in pilots
External advisors and consultants
Regulatory bodies for clarification and guidance
Revision History Summary
Major Revisions:
v1.0 (January 2026): Initial publication
Complete framework aligned to NIST AI RMF 1.0
GenAI-specific controls from NIST AI 600-1
EU AI Act compliance mapping
32 controls across 15 categories
Complete templates and appendices
Future Planned Updates:
Q2 2026: Incorporation of EU AI Act implementing acts
Q3 2026: Enhanced automation and tooling guidance
Q4 2026: Industry-specific addendums (healthcare, financial services)
Q1 2027: Annual review and regulatory alignment update
Quick Reference Card
For AI System Owners - Getting Started:
Submit Intake Form → ai-arb-intake@[organization].com
Calculate Risk Tier → Use rubric in Section 8
Prepare Artifacts → Templates in Section 11
Get Approvals → Per tier requirements in Section 9
Deploy with Monitoring → Follow Section 10 workflow
Risk Tier Quick Reference:
Tier 1 (0-6 points): Basic documentation + Business/Engineering Owner approval
Tier 2 (7-12 points): Full TEVV + ARB + Security/Privacy/Domain approval
Tier 3 (13-18 points or override): Red-team testing + Executive risk acceptance
Emergency Contacts:
AI Governance Help: ai-governance@[organization].com
P1 Incidents: PagerDuty "AI-Incident-P1"
Security Issues: ai-security@[organization].com
Training:
Level 1 (All): 30 min - ai-training@[organization].com
Level 2 (AI Users): 2 hrs - Required for system submission
Level 3 (Reviewers): 1 day - Required for ARB/review roles
END OF AI GOVERNANCE FRAMEWORK - VERSION 1.0



