Chapter 3. WIA Emotion AI Standard Overview

Hongik Ingan (ๅผ˜็›Šไบบ้–“)

"Benefit All Humanity"

The WIA Emotion AI Standard provides a comprehensive framework for ethical, accurate, and interoperable affective computing. This chapter offers a single-screen overview of the four-phase architecture, the classification framework, the multimodal fusion strategies, and the design principles, so that decision-makers can grasp the standard's structure and value before turning to the detailed Phase 1โ€“4 specifications in Chapters 4 through 8. It also defines the simulator-to-volume mapping table (ยง3.5b) that anchors English-edition vocabulary to the simulator at https://wiastandards.com/emotion-ai/simulator/.

3.1 Mission and Goals

3.1.1 Mission Statement

The WIA Emotion AI Standard sets out to establish a universal, open framework for emotion-recognition systems that enables innovation and interoperability while putting human well-being, privacy, and accuracy first. The mission is not merely the definition of a technical specification; it pursues four axes simultaneously โ€” protection of user rights, fair competition in the industry ecosystem, consensus among academia, industry, and government, and compatibility across the major regulatory regions (Korea, EU, United States, and the wider Asia-Pacific). The four axes are mutually supportive; standards that emphasise only one of them have historically failed either by under-adoption or by loss of social trust.

The standard does not seek to replace the work of ISO/IEC, IEEE, W3C, ITU-T, or national-association standards bodies; it positions itself as an interoperability layer over them. Its operating principle is neutrality: it represents no single firm, country, or school of thought. The "one country, one vote" governance rule is adopted to forestall standard capture by large-share firms. This contrasts with sectoral-association models in which voting weight scales with contribution.

3.1.2 Core Goals

Table 3-1. The five core goals of the WIA standard and the benefits they enable
GoalDescriptionBenefit
InteroperabilityCommon data format and APIFreedom from vendor lock-in
AccuracyMinimum accuracy thresholdsTrustworthy results
EthicsPrivacy and consent requirementsResponsible AI
FairnessMandatory bias testingEquitable performance
TransparencyClear documentationUser-comprehensible outputs

These five goals are not mere slogans; they are realised as twenty-four measurable conformance items. Each item has four constituents โ€” a test procedure, a pass criterion, a reporting form, and a corrective-action workflow โ€” so that abstract values become concretely auditable. The twenty-four items are distributed across the five goals (six, five, five, four, four), reflecting the natural decomposition rather than a relative-importance weighting.

3.2 The Four-Phase Architecture

The WIA Emotion AI Standard organises the affective computing stack into four phases, each of which addresses a distinct layer. Each phase can be implemented and certified independently; full value, however, is realised by adopting all four. Phase decomposition deliberately lowers the entry barrier and enables incremental scaling: Phases 1 and 2 may be adopted first; Phase 3 added when real-time processing is needed; Phase 4 introduced when domain-specific regulatory alignment becomes necessary.

Figure 3-1. The four-phase architecture of the WIA Emotion AI Standard
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                    Phase 4: Integration                      โ”‚
โ”‚   Healthcare โ”‚ Education โ”‚ Marketing โ”‚ Automotive โ”‚ XR     โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚                    Phase 3: Streaming Protocol               โ”‚
โ”‚       WebSocket โ”‚ REST โ”‚ Real-time streaming โ”‚ Security    โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚                    Phase 2: API Interface                    โ”‚
โ”‚  Face โ”‚ Voice โ”‚ Text โ”‚ Biosignal โ”‚ Multimodal fusion       โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚                    Phase 1: Data Format                      โ”‚
โ”‚  JSON Schema โ”‚ Emotion โ”‚ AU codes โ”‚ V-A โ”‚ Metadata         โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

3.2.1 Phase 1 โ€” Emotion Data Format

  • Purpose. Standardises the representation of emotion data.
  • Scope. JSON Schema, emotion labels, AU codes, dimensional axes.
  • Detail. Chapter 4.

Phase 1 is the common vocabulary on which every later phase depends and is the most basic certification step. A system that implements only Phase 1 can already qualify as "WIA Compliant"; this lower bar is intentional, lowering the entry barrier to the standard.

3.2.2 Phase 2 โ€” API Interface

  • Purpose. Defines common API endpoints and methods.
  • Scope. Per-modality REST APIs, request and response formats.
  • Detail. Chapter 5.

Phase 2 is the standard transport for Phase 1 data. It follows REST principles and provides per-modality endpoints, so a system that only needs facial-expression analysis is not required to implement voice or text endpoints in order to be conformant. Authentication is supported via OAuth 2.0 and API Key on equal footing.

3.2.3 Phase 3 โ€” Streaming Protocol

  • Purpose. Enables real-time emotion streaming.
  • Scope. WebSocket protocol, frame rate, security.
  • Detail. Chapter 6.

Phase 3 is used by applications that demand real-time processing โ€” driver monitoring, contact-centre call analysis, interactive games. Non-real-time uses (for example, retrospective text-diary analysis) can rely on Phase 2 alone. WebSocket is the recommended primary protocol; a gRPC-streaming adapter is defined in Annex H for microservice environments.

3.2.4 Phase 4 โ€” Integration

  • Purpose. Domain-specific integration guidance.
  • Scope. Healthcare, education, marketing, automotive, XR.
  • Detail. Chapter 7.

Phase 4 reconciles domain-specific regulatory requirements and ethical guidance. Region-specific annexes (Annex KR, Annex EU, Annex US) document how the four-phase architecture meets each regulatory environment.

3.3 Emotion Classification Framework

3.3.1 The Discrete Model โ€” Ekman

The WIA standard adopts seven labels โ€” Ekman's six basic emotions plus Neutral โ€” as the core interoperability layer. The discrete model is intuitive for end-users and lends itself to statistical analysis; it is particularly well-suited to applications such as content recommendation and marketing-effectiveness measurement that compare label distributions across users.

Table 3-2. The seven core emotion labels of the WIA standard with typical Valence-Arousal ranges
EmotionEnglish labelEmojiTypical V-A range
Happinesshappiness๐Ÿ˜ŠV: 0.5 to 1.0; A: 0.2 to 0.8
Sadnesssadness๐Ÿ˜ขV: โˆ’0.8 to โˆ’0.3; A: โˆ’0.5 to 0.1
Angeranger๐Ÿ˜ V: โˆ’0.7 to โˆ’0.2; A: 0.3 to 0.9
Fearfear๐Ÿ˜จV: โˆ’0.7 to โˆ’0.2; A: 0.4 to 0.9
Disgustdisgust๐ŸคขV: โˆ’0.8 to โˆ’0.3; A: โˆ’0.1 to 0.5
Surprisesurprise๐Ÿ˜ฎV: โˆ’0.2 to 0.5; A: 0.5 to 1.0
Neutralneutral๐Ÿ˜V: โˆ’0.2 to 0.2; A: โˆ’0.2 to 0.2

Display labels in the user's language are layered over a fixed lower-case English key (happiness, sadness, โ€ฆ) so that locale-specific display labels in Korean, Japanese, Chinese, or Arabic resolve to the same internal label key. This separation is compatible with the W3C EmotionML 1.0 vocabulary mechanism (W3C, 2014).[1]

3.3.2 The Dimensional Model โ€” Valence-Arousal

The WIA standard supports the dimensional model in parallel with the discrete model. The dimensional model composes naturally with regression-style machine-learning and is strong on representing subtle change and mixed emotion. Music recommendation, digital-meditation applications, and adaptive-difficulty games โ€” all of which need continuous tracking of user state โ€” derive particular value from the dimensional output.

Figure 3-2. The Valence-Arousal space
                    +1.0 (high arousal)
                          โ”‚
                   Anger  โ”‚  Excited
                          โ”‚
    -1.0 โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ +1.0
    (negative)            โ”‚           (positive)
                          โ”‚
                   Sadnessโ”‚  Calm
                          โ”‚
                    -1.0 (low arousal)

Range:
  Valence: -1.0 (most negative) to +1.0 (most positive)
  Arousal: -1.0 (low energy) to +1.0 (high energy)

WIA Phase 1 recommends populating both the discrete and dimensional outputs simultaneously: discrete labels feed user-facing UIs, while V-A coordinates support statistical analysis and time-series tracking. Conversion between the two representations is straightforward in either direction using the typical ranges of Table 3-2. When the dimensional output's nearest discrete-label V-A centre exceeds a configurable distance threshold, the UI is recommended to fall back to a "mixed" or "uncertain" label rather than coerce a single category โ€” a concrete instance of the standard's transparency principle.

3.3.3 Extended Emotion Labels

Table 3-3. WIA extended emotion labels โ€” five categories
CategoryExtended labels
Positive, high arousalexcited, elated, enthusiastic, amused
Positive, low arousalcontent, relaxed, calm, serene
Negative, high arousalstressed, anxious, frustrated, irritated
Negative, low arousalbored, tired, depressed, melancholic
Cognitive statesconfused, focused, interested, engaged

Culturally specific affect labels โ€” for example culturally specific forms of grief, longing, embarrassment, or sustained-relational attachment โ€” may be added through the user-defined-label mechanism. A user-defined label must declare an English key, a display name, a typical V-A range, and a citation; once sufficiently adopted, the WIA Steering Committee can promote it to an official label in the next major version.

3.4 FACS Integration

3.4.1 Supported Action Units

The WIA standard supports all forty-four Action Units defined by FACS (Ekman & Friesen, 1978; Ekman, Friesen & Hager, 2002).[2]

Table 3-4. Action Units supported by the WIA standard, organised by region
AU rangeRegionCount
AU1โ€“AU7Upper face โ€” brows, forehead7
AU9โ€“AU17Nose, upper lip8
AU18โ€“AU28Lower face โ€” lips, chin11
AU41โ€“AU46Eyelids6
AU51โ€“AU58Head pose8
AU61โ€“AU64Eye position4

The sixteen mandatory core AUs are listed in Table 1-5; the remaining twenty-eight are optional. Medical-and research-grade applications are advised to support all forty-four; consumer applications are typically adequate with the sixteen core. The industry reference for AU-recognition accuracy is agreement with OpenFace 2.0 (Baltruลกaitis et al., 2018), which the WIA conformance suite uses as a partial benchmark.[3]

3.4.2 AU Intensity Coding

Table 3-5. WIA AU intensity coding โ€” mapping to the FACS five-grade scale
Intensity valueFACS gradeMeaning
0.0โ€”Not active
0.01โ€“0.20A โ€” TraceTrace
0.21โ€“0.40B โ€” SlightSlight
0.41โ€“0.60C โ€” MarkedMarked
0.61โ€“0.80D โ€” PronouncedPronounced
0.81โ€“1.00E โ€” MaximumMaximum

Conversion between the floating-point and five-grade representations is lossless; conformant systems must emit at least one of the two. AU output is the primary means by which a face-analysis system exposes its internal evidence to auditors, supporting the EU AI Act's transparency and human-oversight requirements.

3.5 Supported Modalities

3.5.1 Facial Expression Analysis

Table 3-6. WIA facial-expression modality โ€” input, resolution, frame-rate, output, latency
ItemSpecification
Input typeImage (JPEG, PNG) or video (H.264, VP9)
Resolutionโ‰ฅ 480p; recommended โ‰ฅ 720p
Frame rateโ‰ฅ 15 fps; recommended 30 fps
OutputEmotion label, AU intensities, V-A coordinates
Latency target< 100 ms per frame

3.5.2 Voice and Speech Analysis

Table 3-7. WIA voice modality โ€” input format, sample rate, channels, features, output
ItemSpecification
Input typeAudio (WAV, MP3, WebM)
Sample rateโ‰ฅ 16 kHz; recommended 44.1 kHz
ChannelsMono or stereo
FeaturesPitch, intensity, speech rate, voice quality, prosody
OutputEmotion label, V-A coordinates, confidence

3.5.3 Text Sentiment Analysis

Table 3-8. WIA text modality โ€” input, languages, length, output, features
ItemSpecification
Input typeUTF-8 text
Languages supportedโ‰ฅ 100 languages
Maximum length10,000 characters per request
OutputSentiment polarity, emotion label, per-entity emotion
FeaturesSarcasm detection, aspect-based sentiment, intensity

3.5.4 Biosignal Analysis

Table 3-9. WIA biosignal modality โ€” signals, sample rates, format, output
ItemSpecification
Supported signalsECG / HR, EDA / GSR, EEG, respiration
Sample rateHR โ‰ฅ 1 Hz; EDA โ‰ฅ 4 Hz; EEG โ‰ฅ 128 Hz
FormatJSON array or CSV
OutputArousal, stress index, engagement

Each modality has distinct strengths and weaknesses, and combinations are recommended. Facial expression is strong for discrete-emotion classification; voice is strong for arousal estimation; text contributes intent and context; biosignal is hardest to fake. These complementary properties motivate the multimodal-fusion strategies of ยง3.6. In medical-grade applications all four modalities are recommended; clinical decisions based on a single modality are excluded from conformance.

3.5b Simulator Five-Panel ENUM and Provisional-Threshold Mapping

The interactive simulator that ships with this volume โ€” accessible at https://wiastandards.com/emotion-ai/simulator/ โ€” is organised as five working panels. Each panel maps one-to-one onto a phase of the four-phase architecture (Phase 1 to Phase 4). This section consolidates the panel headers, core ENUMs, and provisional thresholds used in the simulator so that the volume's vocabulary remains aligned with the simulator across editions. Where vocabulary diverges, the ENUM as displayed in the simulator takes precedence.

Table 3-5b. WIA Simulator 5-Panel ENUM / Provisional-Threshold Mapping
PanelSimulator headerCore ENUMProvisional thresholdVolume mapping
0๐Ÿ“Š Emotion DataINPUT, OPTION, TEXTAREA, Subject ID, Primary Emotion, Confidence (0โ€“1), Valence (โˆ’1 to +1), Arousal (โˆ’1 to +1), Action Units (FACS), ModalityPer-modality classification accuracy: provisional 76.8% to 92.3%Phase 1 โ€” Data format (Chapter 4)
1๐Ÿ”ข AnalysisValence-Arousal four-quadrant, High Arousal ยท Low Arousal ยท Negative ยท PositiveAnalysis window: provisional 0.2 s to 0.3 sChapter 1 ยง1.3.2 โ€” Dimensional model
2๐Ÿ“ก ProtocolSTREAMING, STOPPED, message format (JSON, MessagePack, Protocol Buffers)Streaming latency: provisional โ‰ค 300 ms (see ยง6.3 for measured values)Phase 3 โ€” Streaming protocol (Chapter 6)
3๐Ÿ”— IntegrationExternal-system adapters, domain-specific standards (FHIR, SCORM, etc.)(No simulator-defined threshold โ€” see ยง7.3, ยง7.4 for domain thresholds)Phase 4 โ€” Integration (Chapter 7)
4๐Ÿงช Emotion TestPredicted Emotion output, six basic emotions (Ekman) with dimensional model auxiliaryClassification response time: provisional 0.2 s to 0.3 sChapter 8 โ€” Implementation and certification (test procedure)

The mapping is the alignment reference between the volume's ENUMs and thresholds and those of the simulator; both are updated in lockstep. Where the volume cites a threshold not explicitly stated in the simulator, the citation is rendered in provisional form ("provisional", "approximately", "expected") so that the simulator's authority is preserved.

3.6 Multimodal Fusion

3.6.1 Fusion Strategies

Table 3-10. Four multimodal-fusion strategies and their use cases
StrategyDescriptionUse case
Early fusionCombine raw features before classificationModalities synchronised in time
Late fusionCombine classification outputsModalities are independent
Decision fusionVoting or weighted average over decisionsSimple, robust approach
Attention fusionContext-conditioned learned weightsModality reliability varies over time

Late fusion and decision fusion are the most common in industrial systems because they preserve modular development; early fusion and attention fusion can deliver higher accuracy but at the cost of more complex training. Attention fusion has gained traction with the rise of transformer-based multimodal models and is particularly effective in environments where modality reliability fluctuates over time (call-noise variation, in-cabin lighting variation, and the like). The choice of strategy also depends on the legal and consent basis: voice analysis on telephony channels typically requires two-party consent under telecommunications-secrecy law, whereas text analysis on user-supplied input may not, so the legal "processing units" of two modalities can differ.

3.6.2 Modality Weighting

Figure 3-3. Default multimodal weights โ€” face 0.40 ยท voice 0.25 ยท text 0.20 ยท biosignal 0.15
Default weights (adjustable):
  Face:      0.40 (highest reliability for discrete-emotion classification)
  Voice:     0.25 (strong for arousal estimation)
  Text:      0.20 (context-dependent)
  Biosignal: 0.15 (hard to fake but noisy)

Adjustment criteria:
  - Signal quality
  - Context (e.g. voice-only call)
  - Cultural factors
  - Per-user calibration

3.7 Design Principles

3.7.1 Core Principles

  1. Privacy by Design. Minimal data collection; consent obligation.
  2. Transparency. Clear disclosure that emotion AI is in use.
  3. Accuracy. Minimum thresholds together with demographic fairness.
  4. Interoperability. Standard formats enable data portability.
  5. Extensibility. User-defined emotions and modalities are supported.
  6. Cultural sensitivity. Cultural difference is reflected in design.
  7. Human oversight. Human review of automated decisions is possible.

3.7.2 Technical Principles

  1. JSON-based. Human-readable; widely supported.
  2. Semantic versioning. Clear upgrade paths (SemVer 2.0).
  3. REST and WebSocket. Standard web protocols.
  4. Confidence scores. Always include an uncertainty estimate.
  5. Timestamps. Time-series analysis is enabled.
  6. UTF-8. Equal treatment of every language, including Korean.
  7. Minimal dependencies. Standard implementation must not be tied to a specific cloud vendor.

3.8 Certification Levels

Table 3-11. The three certification levels of the WIA standard
LevelNameRequirementsUse case
1CompliantData-format conformance; 75% accuracyResearch, prototypes
2CertifiedFull API conformance; 80% accuracy; bias testingCommercial products
3Certified PlusAll requirements; 85% accuracy; external auditHealthcare; sensitive applications

Level transitions require only the additional tests for the next level rather than re-certification of prior tests. Certificates are valid for twenty-four months; renewal triggers either partial or full re-testing depending on whether the model, code, or training data have changed. A renewal-due notification is issued thirty days before expiry; failure to begin renewal causes automatic expiry, after which the full certification process must be re-run.

3.9 Note on Korean Edition Content

The Korean edition of this volume includes additional sections specific to Korea: detailed alignment with national-association standards (industrial standards, ICT-association standards, and a national cyber-security agency self-assessment toolkit), a five-scenario industry-application section grounded in named domestic enterprises, a five-stage adoption roadmap calibrated to typical Korean enterprise timelines (six to twelve months), and a governance section describing how Korean institutions can contribute to standards revisions through formal channels.

This English edition deliberately abstracts those passages. References to specific named Korean industrial standards, association standards, public-agency tools, and enterprise scenarios become "leading domestic standards bodies", "the relevant national cyber-security agency self-assessment toolkit", "leading domestic conglomerates", "leading domestic universities, government agencies, telecom operators, and platform companies", and "leading commercial SDK vendors". The conformance requirements themselves are identical between the two editions.

3.10 Chapter Summary

Nine key takeaways.

  1. Four-phase architecture. Data format โ†’ API โ†’ protocol โ†’ integration.
  2. Dual model. Both discrete (Ekman) and dimensional (V-A) representations are supported.
  3. FACS. All forty-four AUs are codable.
  4. Four modalities. Face, voice, text, biosignal.
  5. Multimodal fusion. Four strategies are supported.
  6. Three certification levels. Compliant, Certified, Certified Plus.
  7. Simulator alignment. Table 3-5b maps the volume to the simulator's five panels.
  8. Incremental adoption. A phased adoption path lowers the barrier for newcomers.
  9. Governance transparency. All revision proposals and votes are recorded openly in the public GitHub repository.

3.11 Review Questions

  1. Draw the four-phase architecture (data format, API, protocol, integration) and describe each phase's role.
  2. Compare the discrete and dimensional emotion models and explain why simultaneous output is recommended.
  3. Tabulate the seven core emotion labels of the WIA standard with their typical Valence-Arousal ranges.
  4. Describe the user-defined-label mechanism and the four metadata items it requires.
  5. Explain the four multimodal-fusion strategies (early, late, decision, attention) with at least one industrial example each.
  6. Describe the three certification levels with their use cases and accuracy thresholds.
  7. Describe the role of Table 3-5b in maintaining alignment between the volume and the simulator.
  8. Explain how the standard treats culturally specific affect labels through the user-defined-label mechanism.

3.12 Looking Ahead

Chapter 4 turns to Phase 1 โ€” Emotion Data Format โ€” and treats the JSON Schema, field specifications, and worked examples in depth. Phase 1 is the common vocabulary on which every later phase depends and is the longest chapter in the volume. After Chapter 4 the reader will have enough understanding to write an adapter that converts an arbitrary emotion-recognition system's output into the WIA Phase 1 format. The standard's evolution roadmap is recorded in the public GitHub repository.[99]

Chapter 3 Endnotes

  1. W3C. (2014). Emotion Markup Language (EmotionML) 1.0, W3C Recommendation, 22 May 2014. https://www.w3.org/TR/emotionml/. โ†‘
  2. Ekman, P., & Friesen, W. V. (1978). Facial Action Coding System: A Technique for the Measurement of Facial Movement. Consulting Psychologists Press, ISBN 0-931835-01-1. Updated edition: Ekman, Friesen & Hager (2002). โ†‘
  3. Baltruลกaitis, T., Zadeh, A., Lim, Y. C., & Morency, L.-P. (2018). OpenFace 2.0: Facial behavior analysis toolkit. In Proceedings of the 13th IEEE International Conference on Automatic Face & Gesture Recognition, 59โ€“66. DOI 10.1109/FG.2018.00019. โ†‘
  4. ISO/IEC 22989:2022. Information technology โ€” Artificial intelligence โ€” Artificial intelligence concepts and terminology. https://www.iso.org/standard/74296.html.
  5. ISO/IEC 23053:2022. Framework for Artificial Intelligence (AI) Systems Using Machine Learning (ML). https://www.iso.org/standard/74438.html.
  6. ISO/IEC 23894:2023. Information technology โ€” Artificial intelligence โ€” Guidance on risk management.
  7. ITU-T. (2018). Recommendation F.748.11: Emotion-aware multimedia services.
  8. IEEE 7000-2021. IEEE Standard Model Process for Addressing Ethical Concerns During System Design. DOI 10.1109/IEEESTD.2021.9536679.
  9. IEEE 7003-2024. IEEE Standard for Algorithmic Bias Considerations.
  10. NIST. (2023). AI Risk Management Framework (AI RMF) 1.0. NIST AI 100-1. DOI 10.6028/NIST.AI.100-1.
  11. Russell, J. A. (1980). A circumplex model of affect. Journal of Personality and Social Psychology 39(6), 1161โ€“1178. DOI 10.1037/h0077714.
  12. Plutchik, R. (1980). A general psychoevolutionary theory of emotion. In Theories of Emotion. Academic Press, 3โ€“33. DOI 10.1016/B978-0-12-558701-3.50007-7.
  13. Baltruลกaitis, T., Ahuja, C., & Morency, L.-P. (2019). Multimodal Machine Learning: A Survey and Taxonomy. IEEE Transactions on Pattern Analysis and Machine Intelligence 41(2), 423โ€“443. DOI 10.1109/TPAMI.2018.2798607.
  14. IETF RFC 6455. (2011). The WebSocket Protocol. DOI 10.17487/RFC6455.
  15. SemVer 2.0.0. Semantic Versioning Specification. https://semver.org/.
  16. WIA Standards public repository (emotion-ai folder), MIT-licensed source for the simulator, specification, API reference, and ebook assets cited throughout this volume: WIA-Official/wia-standards-public/tree/main/emotion-ai. The standard's evolution roadmap, revision history, and SDK source code are maintained openly in this repository, where the WIA standards committee records its formal verification of all primary sources cited in this chapter. โ†‘