Chapter 4: Sign Language Corpus and Annotation Standards (ELAN/SignBank/HamNoSys)

Sign-language corpora are the foundational asset of any sign-recognition system. This chapter covers the international standard annotation frameworks — ELAN, SignBank, HamNoSys, and SignWriting — and Korea's sign-language corpus infrastructure.

4.1 The Nature of Sign-Language Corpora

A sign-language corpus is a collection of sign-language video together with attached linguistic annotation. Unlike text corpora of spoken languages, sign corpora inherently contain video data and therefore require separate standards for storage capacity, transmission bandwidth, and personal-information protection. The WIA Sign Language Recognition Standard standardizes the following four corpus domains.

  1. Video specifications: Resolution, frame rate, color space, compression codec, metadata.
  2. Annotation formats: ELAN, SignBank, HamNoSys, SignWriting compatible.
  3. Lexical identifiers (gloss): A unique ID is assigned to each sign vocabulary item.
  4. Personal-information protection: Face mosaicking, skeleton extraction, and other de-identification procedures.

Corpus metadata format can be validated in the simulator Data Format panel.

4.1.1 Corpus Units

A sign-language corpus is structured into the following five levels.

LevelDescriptionExample
SessionOne recording session of one signerSigner #001, recorded 1 March 2024
RecordingOne continuous videosession_001_rec_05.mp4
SentenceOne sentence-level segment"Hello. Nice to meet you."
SignOne lexical signHELLO, MEET, GLAD
FrameOne video frame1/30 sec at 30 fps

4.2 ELAN — De-facto Standard Annotation Tool

ELAN (EUDICO Linguistic Annotator)1) is a multimedia annotation tool developed by the Max Planck Institute for Psycholinguistics. Since its initial release in 2002 it has become the de facto standard tool for sign-language corpus annotation worldwide. As of 2024 the current version is 6.7.

4.2.1 ELAN Annotation Format

An ELAN annotation file (.eaf) is an XML-based multi-tier structure.

4.2.2 Korean Sign Language ELAN Template

The National Institute of the Korean Language (NIKL) released a "Korean Sign Language ELAN Annotation Template" in 2020. It comprises the following ten tiers.

Tier nameAnnotation contentExample values
gloss-kslKSL lexical IDHELLO, MEET, GLAD
gloss-koKorean gloss안녕하세요
handshape-rightRight-hand handshape codeHS_FLAT, HS_FIST
handshape-leftLeft-hand handshape codeHS_FLAT, HS_INDEX
locationLocation codeLOC_CHEST, LOC_FOREHEAD
movementMovement codeMOV_STRAIGHT, MOV_CIRCULAR
orientationPalm orientationOR_UP, OR_FORWARD
nmm-eyebrowEyebrow NMMBROW_RAISE, BROW_FROWN
nmm-mouthMouth NMMMOUTH_OPEN, MOUTH_PONG
nmm-headHead movementHEAD_NOD, HEAD_SHAKE

The WIA standard adopts the NIKL ELAN template as the primary KSL annotation format.

4.3 SignBank — Sign Dictionary Standard

SignBank2) is an open-source platform for sign-language dictionaries of ASL, BSL, NGT (Sign Language of the Netherlands), Auslan (Australian Sign Language), and others. It is operated by Radboud University. Each dictionary contains:

4.3.1 KSL Dictionary Alignment with SignBank

The NIKL KSL Dictionary introduced SignBank-compatible JSON export in 2023. To ensure international interoperability of KSL Dictionary data, the WIA standard adopts the SignBank JSON format.

4.4 HamNoSys — Phonological Feature Notation

HamNoSys (Hamburg Notation System)3), developed by the University of Hamburg in 1985, encodes sign-language signals using approximately 200 symbols that represent the five distinctive features (handshape, location, movement, orientation, NMM). HamNoSys symbols are Unicode-registered, allowing text-based processing, search, and database storage.

4.4.1 HamNoSys Example

HamNoSys notation for the Korean Sign Language SCHOOL sign:

HSꜜ⠁ LOC⠂CHEST MOV⠃CIRCULAR OR⠄FORWARD

This denotes "a flat hand moving in a circular trajectory in front of the chest with palm facing forward." HamNoSys is widely used in dictionaries, grammar books, and research papers.

4.5 SignWriting — Visual Sign Notation

SignWriting4), developed by Valerie Sutton in 1974, is a visual notation system for sign languages. ISWA 2010 (International SignWriting Alphabet 2010) defines approximately 600 symbols. It is used in print, education, and literary publication.

4.5.1 SignWriting vs. HamNoSys

SignWriting represents "visual form," whereas HamNoSys decomposes "phonological features." SignWriting is generally regarded as more Deaf-friendly within the Deaf community, while HamNoSys is preferred for precision linguistic analysis.

4.6 Korean Sign Language Corpus Infrastructure

4.6.1 NIA AI Hub "AI Training Korean Sign Language Video"

The NIA AI Hub Korean Sign Language Video dataset5) has the following specifications.

ItemSpecification
Number of signersApproximately 1,200
Vocabulary sizeApproximately 110,000
Total video length1,000 hours
RGB resolution1920×1080 @ 30 fps
Depth resolution640×480 @ 30 fps (Azure Kinect)
Annotation toolsELAN + in-house tool
NMM annotationMajor NMMs including interrogative, negation, emphasis
Public licenseAvailable for research and commercial use after AI Hub registration

4.6.2 NIKL Korean Sign Language Corpus

NIKL has operated the "Korean Sign Language Corpus Construction Project" since 2014. As of 2024 it has accumulated approximately 4,000 lexical entries, 50 hours of video, and 200 signers. The corpus is publicly released for research purposes and complements the NIA AI Hub dataset.

4.6.3 KETI Sign Language Dataset

Released by the Korea Electronics Technology Institute (KETI) in 2019, the dataset comprises 419 lexical entries, approximately 14 hours of video, and 14 signers. It is the standard benchmark for isolated SLR.

4.6.4 KAIST KSL Hello

The KAIST AI Graduate School has released a dataset of 200 KSL greetings and daily expressions, comprising 10 signers and approximately 8 hours of video including Azure Kinect depth data.

4.7 Gloss ID Assignment Rules

A gloss is an uppercase-letter code that represents a sign-language sign in text. The WIA standard adopts the following rules.

KSL vocabulary entries record both English and Korean glosses. Combined with ISO 639-3 language codes, the standard recommends the global unique identifier format "kvk:SCHOOL."

4.8 Personal-Information Protection — De-Identification Procedures

Because sign-language video reveals the signer's face, it constitutes "sensitive information" under Article 2 of the Personal Information Protection Act. The standard recommends the following de-identification procedures.

4.8.1 Face Mosaicking

Mosaicking is applied to the signer's face region. However, when NMM recognition is required, the facial information is lost and the technique is unsuitable.

4.8.2 Skeleton Extraction with Video Discard

After extracting 543 keypoints via MediaPipe Holistic, the original video is discarded and only the keypoints are retained. This is the approach recommended by the KISA "Sign Language Video Personal-Information De-identification Guideline" (2023).6)

4.8.3 Avatar Synthesis

After extracting the original signer's skeleton, the same sign is synthesized through a virtual avatar before release. This fully protects the signer's identity while preserving NMMs.

4.8.4 Consent Procedure

Explicit consent for recording, publication, and research use must be obtained from the signer in Korean Sign Language, in line with the Korean Sign Language Act and Article 30 of the Anti-Discrimination Against and Remedies for Persons with Disabilities Act. Consent obtained through hearing-style printed forms does not guarantee sufficient understanding by Deaf signers; "non-discriminatory consent" must be ensured.

4.9 Korean Corpus Infrastructure Mapping

4.99 Closing Remarks and Korean Infrastructure Cross-Reference

The technical specifications described in this chapter all serve the broader goal of guaranteeing the everyday communication rights of the Deaf community. Korean Sign Language is, under Article 2 of the Korean Sign Language Act, a public language with status equal to Korean, and the WIA Sign Language Recognition Standard provides the technical underpinning for this legal status.

The standard operates on the Korean national infrastructure: the Korean Association of the Deaf (KAD), the National Institute of the Korean Language (NIKL), the National Information Society Agency (NIA), the Electronics and Telecommunications Research Institute (ETRI), the Korea Advanced Institute of Science and Technology (KAIST), the Korea Institute of Science and Technology Information (KISTI), the Telecommunications Technology Association (TTA), the Korean Standards Association (KSA), the Korean Agency for Technology and Standards (KATS), the Korea Internet and Security Agency (KISA), the Korea Laboratory Accreditation Scheme (KOLAS), and the Korea Communications Agency (KCA), as well as government ministries including the Ministry of Culture, Sports and Tourism (MCST), the Ministry of Health and Welfare (MOHW), the Ministry of Education (MOE), the National Institute of Special Education (NISE), the Korea Communications Commission (KCC), the Ministry of Science and ICT (MSIT), the National Human Rights Commission of Korea (NHRCK), the Ministry of Employment and Labor (MOEL), the Ministry of the Interior and Safety (MOIS), and the Ministry of Justice.

Broadcasters KBS, MBC, SBS, EBS, National Assembly Television, and Arirang International Broadcasting plan to adopt the standard's automatic captioning system; 5G operators SK Telecom, KT, and LG U+ apply the standard to Deaf telecommunication relay services; and AI providers Naver Clova, Kakao i, LG AI Research, Kakao Brain, and SK Telecom X publish KSL recognition APIs compatible with the standard.

The eighteen schools for the Deaf nationwide (Seoul School for the Deaf, Daejeon School for the Deaf, Busan Sungsim School, Incheon Sunhwa School, Gwangju Sunmyeong School, Daegu Yeonghwa School, Gangwon Provincial Dowon School, Chungbuk Cheongju Sungsim School, Chungnam Cheonan Inae School, Jeonbuk Iksan Jeil School, Jeonnam Gwangju Yeonghwa School, Gyeongbuk Andong Yeongmyeong School, Gyeongnam Jinju Hyegwang School, Jeju Yeongji School, Incheon Cheonghak School, Gyeonggi Ansan Jahae School, Ulsan Meari School, and Sejong Sarang School) serve as the standard's KSL education hubs.

This standard is published as an open standard under the MIT license; all simulator code, specifications, and example code are openly available at GitHub WIA-Official/wia-standards-public/tree/main/sign-language. Our hope is that this standard enables Deaf people who use Korean Sign Language to communicate more freely, and that this standard becomes the foundation for sign-language recognition standards across the Asia-Pacific region.

Chapter 4 Endnotes

  1. ELAN (EUDICO Linguistic Annotator) Version 6.7, Max Planck Institute for Psycholinguistics, 2024. De-facto standard tool for sign-language corpus annotation.
  2. SignBank Platform, Radboud University, ongoing. Multi-country sign-dictionary platform for ASL, BSL, NGT, and others.
  3. Siegmund Prillwitz et al., HamNoSys: Hamburg Notation System for Sign Languages, University of Hamburg, 1989, revised 4.0 in 2004. Phonological feature notation system.
  4. Valerie Sutton, International SignWriting Alphabet 2010 (ISWA 2010), Center for Sutton Movement Writing, 2010. Visual notation system for sign languages.
  5. AI Training Korean Sign Language Video Dataset, National Information Society Agency (NIA) AI Hub, 2021–present. 1,000-hour KSL corpus.
  6. Sign Language Video Personal-Information De-identification Guideline, Korea Internet and Security Agency (KISA), 2023.
  7. "Korean Sign Language ELAN Annotation Template," National Institute of the Korean Language, 2020. Standard Korean-language template for KSL corpus annotation.
  8. TTAK.KO-10.1156, Korean Sign Language Data Annotation Standard, Telecommunications Technology Association (TTA), 2022.
  9. WIA Standards Public Repository (sign-language folder), MIT License, GitHub: WIA-Official/wia-standards-public/tree/main/sign-language — Reference implementations of the ELAN template, SignBank export, HamNoSys parser, and SignWriting renderer defined in this chapter are openly published in this repository.