YPAI
Services Data Industries Company
AI Data & Evaluation
Data collection and sourcing Consent-led multimodal collection. Dataset licensing Rights-cleared datasets, ready to license. Annotation and curation Labelling, review and adjudication. Model and agent evaluation Human evaluation and regression testing. Explore AI Data & Evaluation Create, source and evaluate the data your AI depends on.
AI Implementation
Discovery and architecture Scope the use case and the system design. RAG and knowledge systems Retrieval over your own knowledge. Agents and workflow automation Agents and automation in production. Private and enterprise deployment Private, controlled deployment. Explore AI Implementation Turn a defined AI use case into a system you can operate.
Delivery
Connected Delivery Data, evaluation and implementation under one structure. Pilots Validate the delivery method before scale.
Explore all services
AI Data & Evaluation
Speech & Audio Data Multilingual speech, acoustic environments and voice data. Image, 3D & Sensor Data Images, documents, multi-view data, LiDAR and sensor fusion. Video, Physical AI & Robotics Data On-camera, conversational, egocentric and robotics data. Dataset Licensing & Sourcing Rights-cleared datasets, bespoke sourcing and acquisition. Annotation & Data Production Ontology design, labelling, review and model-ready delivery. Model & Agent Evaluation Human evaluation, multilingual testing and failure analysis.
Explore AI Data & Evaluation
Operating conditions
AI Companies & Model Developers Training data, preference data and evaluation loops. Automotive & Mobility In-cabin speech, perception, video and sensor data. Financial Services Document AI, knowledge systems and traceability. Healthcare & Life Sciences Specialist data, domain review and privacy-sensitive work. Industrial & Energy Field data, operational workflows and integration. Public Sector Controlled data operations and reviewable AI systems.
Explore industry solutions
Company
About YPAI Company, mission, operating model and delivery history. Partnerships Commercial, technology and delivery collaboration. AI Blog Research, technical perspectives and company updates. Contact Projects, partnerships, procurement and general enquiries.
Become a Contributor Contact us
YPAI
AI Data & Evaluation
Data collection and sourcing Dataset licensing Annotation and curation Model and agent evaluation Explore AI Data & Evaluation
AI Implementation
Discovery and architecture RAG and knowledge systems Agents and workflow automation Private and enterprise deployment Explore AI Implementation
Delivery
Connected Delivery Pilots Explore all services
AI Data & Evaluation
Speech & Audio Data Image, 3D & Sensor Data Video, Physical AI & Robotics Data Dataset Licensing & Sourcing Annotation & Data Production Model & Agent Evaluation Explore AI Data & Evaluation
Operating conditions
AI Companies & Model Developers Automotive & Mobility Financial Services Healthcare & Life Sciences Industrial & Energy Public Sector Explore industry solutions
About YPAI Partnerships AI Blog Contact
Contact us Become a Contributor

Speech data

AI Act Risk Classification and Training Data

Last updated: July 2026

The EU AI Act (Regulation 2024/1689) classifies AI systems into risk categories. Classification determines regulatory obligations, including requirements for training data governance under Article 10.

On this page

  • 1. The four risk categories
  • 2. How high-risk classification works
  • 3. Training data requirements
  • 4. Procurement implications
  • 5. YPAI role in AI Act compliance
  • 6. Classification responsibility
  • 7. Key dates
  • 8. Official sources

1. The four risk categories

The AI Act defines four levels of risk. Regulatory obligations increase with risk level.

Unacceptable risk: prohibited

AI systems that pose clear threats to safety, livelihoods, or fundamental rights. These are banned outright. Examples include social scoring systems, real-time biometric identification in public spaces (with limited exceptions), and AI that exploits vulnerabilities or uses subliminal manipulation. Training data is not relevant for prohibited systems: they cannot be deployed.

High risk: regulated

AI systems in safety-critical applications or sensitive use cases. These require conformity assessment, technical documentation, and ongoing compliance obligations including training data governance. Training data requirements under Article 10 apply.

Limited risk: transparency obligations

AI systems that interact directly with people without significant implications. Requirements focus on disclosure, ensuring users know they are interacting with AI. Examples include chatbots and AI-generated content. Training data governance is not mandated, though good practice applies.

Minimal risk: unregulated

AI systems with negligible or no impact on rights or safety. No mandatory requirements. Examples include spam filters and AI-enabled games.

2. How high-risk classification works

An AI system is classified as high-risk through two pathways defined in Article 6.

Pathway 1: safety component (Annex I)

The AI system is a safety component of a product, or is itself a product, covered by EU harmonization legislation listed in Annex I: medical devices, automotive systems, aviation equipment, machinery, lifts, radio equipment, and toys (where safety-relevant). If the product requires third-party conformity assessment and the AI is integral to safety, it is high-risk.

Pathway 2: high-risk use cases (Annex III)

The AI system falls within one of eight areas defined in Annex III:

  • Biometrics
  • Critical infrastructure
  • Education and vocational training
  • Employment
  • Essential services
  • Law enforcement
  • Migration and border control
  • Justice and democratic processes

Automatic high-risk classification

Any AI system that profiles natural persons (automated processing of personal data to evaluate or predict aspects of a person's life, work, health, preferences, or behavior) is always classified as high-risk, regardless of exemptions.

Exemptions from high-risk (limited)

AI systems in Annex III areas may be exempt if they:

  • Perform narrow procedural tasks
  • Improve results of previously completed human activity
  • Detect decision-making patterns without replacing human judgment
  • Perform preparatory tasks only

These exemptions do not apply if the system profiles individuals.

3. Training data requirements for high-risk systems

Article 10 of the AI Act establishes data governance requirements for high-risk AI systems. These requirements apply to training, validation, and testing datasets.

Quality requirements (Article 10.3)

  • Relevant to intended purpose
  • Sufficiently representative
  • Free of errors to best extent possible
  • Complete for intended purpose
  • Statistically appropriate for populations

Governance requirements (Article 10.2)

  • Design choices and collection processes
  • Data preparation (annotation, labeling)
  • Formulation of assumptions
  • Assessment of availability and suitability
  • Examination for biases
  • Identification of data gaps
  • Measures to address identified issues

Contextual requirements (Article 10.4)

Datasets must reflect the specific context of deployment: geographic setting, behavioral context, functional environment, and affected populations.

Documentation requirements: High-risk system providers must maintain technical documentation demonstrating Article 10 compliance. This documentation is subject to review during conformity assessment and may be requested by competent authorities.

4. What this means for training data procurement

Organizations deploying high-risk AI systems must demonstrate that training data meets Article 10 requirements. This obligation sits with the deploying organization, not the data provider. However, training data sourcing decisions directly affect compliance outcomes.

Supports compliance

  • Documented provenance and collection methodology
  • Transparent sampling and representativeness information
  • Bias assessment and limitations disclosure
  • Version control and reproducibility

Creates compliance risk

  • Lack of traceability to data sources
  • No documentation of collection or preparation
  • No assessment of representativeness or gaps
  • No governance artifacts suitable for audit

The question during conformity assessment is not whether training data is available, but whether training data governance can be demonstrated.

5. YPAI's role in AI Act compliance

YPAI is a speech and language data provider. YPAI does not classify AI systems, perform conformity assessments, or issue compliance certifications.

What YPAI provides

  • European speech datasets with documented governance
  • Provenance records and collection methodology
  • Sampling methodology and representativeness information
  • Bias assessment and known limitations disclosure
  • Technical documentation for Article 10 support

What YPAI does not provide

  • System risk classification
  • Conformity assessment services
  • Legal advice on regulatory interpretation
  • Certification of compliance

How this supports customer compliance: Organizations deploying high-risk AI systems can use YPAI's documentation to demonstrate training data governance during conformity assessment. The documentation is structured to address Article 10 requirements.

Whether training data meets the specific requirements for a given AI system depends on the system's intended purpose, deployment context, and affected populations. This assessment is the responsibility of the deploying organization.

6. Classification is the customer's responsibility

YPAI does not determine whether a customer's AI system is high-risk. That determination depends on the system's intended purpose, whether it falls under Annex I or Annex III, whether exemptions apply, and whether the system profiles natural persons.

Organizations uncertain about classification should consult legal counsel or refer to guidance from the European Commission and national competent authorities.

The Commission is required to publish guidelines with practical examples of high-risk and non-high-risk systems by February 2026.

7. Key dates

Date Milestone
August 2024 AI Act entered into force
February 2025 Prohibited practices in effect
August 2025 GPAI model obligations in effect
August 2026 High-risk system obligations in effect (Annex III)
August 2027 High-risk system obligations in effect (Annex I embedded products)

Organizations deploying high-risk AI systems should ensure training data governance is in place before August 2026.

8. Official sources

  • EU AI Act full text: EUR-Lex (Regulation 2024/1689)
  • AI Act Service Desk: ai-act-service-desk.ec.europa.eu
  • European Commission AI policy: digital-strategy.ec.europa.eu
Request AI-Act-ready speech data

Our team can help you navigate AI Act requirements and source compliant speech data for your high-risk AI systems.

Speech data overview · EU AI Act compliant training data · GDPR-compliant speech data · DPA overview

Start with the requirement, not a predefined package.

Bring the objective, current system or dataset, and known operating constraints. YPAI will map the appropriate service line, delivery structure and first validation step.

Contact us Scope a pilot

AI systems, data and evaluation under one accountable delivery model.

New projects · accepting data and AI requirements
Engagement scoped before build
Acceptance defined before delivery
Services
AI Data & Evaluation AI Implementation Controlled Delivery Dataset Licensing
Capabilities
Speech & Audio Image, 3D & Sensor Data Video Data Annotation & Evaluation
Company
About YPAI Partnerships Contact Become a Contributor
Resources & Legal
AI Blog Privacy Terms Cookie Policy Data processing
YPAI · Org. nr. 933 915 778 · Oslo, Norway · Global delivery
Disclaimer LinkedIn ↗ GitHub ↗
EEA-BASED PROCESSING AVAILABLE WHERE REQUIRED · ARTICLE 28 DPA TERMS AVAILABLE
© 2026 YPAI
Install YPAI Faster reopens, offline shell, share-target ready.

Add YPAI to your home screen

Tap the Share button, then Add to Home Screen.