Sign-language corpora are the foundational asset of any sign-recognition system. This chapter covers the international standard annotation frameworks — ELAN, SignBank, HamNoSys, and SignWriting — and Korea's sign-language corpus infrastructure.
A sign-language corpus is a collection of sign-language video together with attached linguistic annotation. Unlike text corpora of spoken languages, sign corpora inherently contain video data and therefore require separate standards for storage capacity, transmission bandwidth, and personal-information protection. The WIA Sign Language Recognition Standard standardizes the following four corpus domains.
Corpus metadata format can be validated in the simulator Data Format panel.
A sign-language corpus is structured into the following five levels.
| Level | Description | Example |
|---|---|---|
| Session | One recording session of one signer | Signer #001, recorded 1 March 2024 |
| Recording | One continuous video | session_001_rec_05.mp4 |
| Sentence | One sentence-level segment | "Hello. Nice to meet you." |
| Sign | One lexical sign | HELLO, MEET, GLAD |
| Frame | One video frame | 1/30 sec at 30 fps |
ELAN (EUDICO Linguistic Annotator)1) is a multimedia annotation tool developed by the Max Planck Institute for Psycholinguistics. Since its initial release in 2002 it has become the de facto standard tool for sign-language corpus annotation worldwide. As of 2024 the current version is 6.7.
An ELAN annotation file (.eaf) is an XML-based multi-tier structure.
The National Institute of the Korean Language (NIKL) released a "Korean Sign Language ELAN Annotation Template" in 2020. It comprises the following ten tiers.
| Tier name | Annotation content | Example values |
|---|---|---|
| gloss-ksl | KSL lexical ID | HELLO, MEET, GLAD |
| gloss-ko | Korean gloss | 안녕하세요 |
| handshape-right | Right-hand handshape code | HS_FLAT, HS_FIST |
| handshape-left | Left-hand handshape code | HS_FLAT, HS_INDEX |
| location | Location code | LOC_CHEST, LOC_FOREHEAD |
| movement | Movement code | MOV_STRAIGHT, MOV_CIRCULAR |
| orientation | Palm orientation | OR_UP, OR_FORWARD |
| nmm-eyebrow | Eyebrow NMM | BROW_RAISE, BROW_FROWN |
| nmm-mouth | Mouth NMM | MOUTH_OPEN, MOUTH_PONG |
| nmm-head | Head movement | HEAD_NOD, HEAD_SHAKE |
The WIA standard adopts the NIKL ELAN template as the primary KSL annotation format.
SignBank2) is an open-source platform for sign-language dictionaries of ASL, BSL, NGT (Sign Language of the Netherlands), Auslan (Australian Sign Language), and others. It is operated by Radboud University. Each dictionary contains:
The NIKL KSL Dictionary introduced SignBank-compatible JSON export in 2023. To ensure international interoperability of KSL Dictionary data, the WIA standard adopts the SignBank JSON format.
HamNoSys (Hamburg Notation System)3), developed by the University of Hamburg in 1985, encodes sign-language signals using approximately 200 symbols that represent the five distinctive features (handshape, location, movement, orientation, NMM). HamNoSys symbols are Unicode-registered, allowing text-based processing, search, and database storage.
HamNoSys notation for the Korean Sign Language SCHOOL sign:
HSꜜ⠁ LOC⠂CHEST MOV⠃CIRCULAR OR⠄FORWARD
This denotes "a flat hand moving in a circular trajectory in front of the chest with palm facing forward." HamNoSys is widely used in dictionaries, grammar books, and research papers.
SignWriting4), developed by Valerie Sutton in 1974, is a visual notation system for sign languages. ISWA 2010 (International SignWriting Alphabet 2010) defines approximately 600 symbols. It is used in print, education, and literary publication.
SignWriting represents "visual form," whereas HamNoSys decomposes "phonological features." SignWriting is generally regarded as more Deaf-friendly within the Deaf community, while HamNoSys is preferred for precision linguistic analysis.
The NIA AI Hub Korean Sign Language Video dataset5) has the following specifications.
| Item | Specification |
|---|---|
| Number of signers | Approximately 1,200 |
| Vocabulary size | Approximately 110,000 |
| Total video length | 1,000 hours |
| RGB resolution | 1920×1080 @ 30 fps |
| Depth resolution | 640×480 @ 30 fps (Azure Kinect) |
| Annotation tools | ELAN + in-house tool |
| NMM annotation | Major NMMs including interrogative, negation, emphasis |
| Public license | Available for research and commercial use after AI Hub registration |
NIKL has operated the "Korean Sign Language Corpus Construction Project" since 2014. As of 2024 it has accumulated approximately 4,000 lexical entries, 50 hours of video, and 200 signers. The corpus is publicly released for research purposes and complements the NIA AI Hub dataset.
Released by the Korea Electronics Technology Institute (KETI) in 2019, the dataset comprises 419 lexical entries, approximately 14 hours of video, and 14 signers. It is the standard benchmark for isolated SLR.
The KAIST AI Graduate School has released a dataset of 200 KSL greetings and daily expressions, comprising 10 signers and approximately 8 hours of video including Azure Kinect depth data.
A gloss is an uppercase-letter code that represents a sign-language sign in text. The WIA standard adopts the following rules.
KSL vocabulary entries record both English and Korean glosses. Combined with ISO 639-3 language codes, the standard recommends the global unique identifier format "kvk:SCHOOL."
Because sign-language video reveals the signer's face, it constitutes "sensitive information" under Article 2 of the Personal Information Protection Act. The standard recommends the following de-identification procedures.
Mosaicking is applied to the signer's face region. However, when NMM recognition is required, the facial information is lost and the technique is unsuitable.
After extracting 543 keypoints via MediaPipe Holistic, the original video is discarded and only the keypoints are retained. This is the approach recommended by the KISA "Sign Language Video Personal-Information De-identification Guideline" (2023).6)
After extracting the original signer's skeleton, the same sign is synthesized through a virtual avatar before release. This fully protects the signer's identity while preserving NMMs.
Explicit consent for recording, publication, and research use must be obtained from the signer in Korean Sign Language, in line with the Korean Sign Language Act and Article 30 of the Anti-Discrimination Against and Remedies for Persons with Disabilities Act. Consent obtained through hearing-style printed forms does not guarantee sufficient understanding by Deaf signers; "non-discriminatory consent" must be ensured.
The technical specifications described in this chapter all serve the broader goal of guaranteeing the everyday communication rights of the Deaf community. Korean Sign Language is, under Article 2 of the Korean Sign Language Act, a public language with status equal to Korean, and the WIA Sign Language Recognition Standard provides the technical underpinning for this legal status.
The standard operates on the Korean national infrastructure: the Korean Association of the Deaf (KAD), the National Institute of the Korean Language (NIKL), the National Information Society Agency (NIA), the Electronics and Telecommunications Research Institute (ETRI), the Korea Advanced Institute of Science and Technology (KAIST), the Korea Institute of Science and Technology Information (KISTI), the Telecommunications Technology Association (TTA), the Korean Standards Association (KSA), the Korean Agency for Technology and Standards (KATS), the Korea Internet and Security Agency (KISA), the Korea Laboratory Accreditation Scheme (KOLAS), and the Korea Communications Agency (KCA), as well as government ministries including the Ministry of Culture, Sports and Tourism (MCST), the Ministry of Health and Welfare (MOHW), the Ministry of Education (MOE), the National Institute of Special Education (NISE), the Korea Communications Commission (KCC), the Ministry of Science and ICT (MSIT), the National Human Rights Commission of Korea (NHRCK), the Ministry of Employment and Labor (MOEL), the Ministry of the Interior and Safety (MOIS), and the Ministry of Justice.
Broadcasters KBS, MBC, SBS, EBS, National Assembly Television, and Arirang International Broadcasting plan to adopt the standard's automatic captioning system; 5G operators SK Telecom, KT, and LG U+ apply the standard to Deaf telecommunication relay services; and AI providers Naver Clova, Kakao i, LG AI Research, Kakao Brain, and SK Telecom X publish KSL recognition APIs compatible with the standard.
The eighteen schools for the Deaf nationwide (Seoul School for the Deaf, Daejeon School for the Deaf, Busan Sungsim School, Incheon Sunhwa School, Gwangju Sunmyeong School, Daegu Yeonghwa School, Gangwon Provincial Dowon School, Chungbuk Cheongju Sungsim School, Chungnam Cheonan Inae School, Jeonbuk Iksan Jeil School, Jeonnam Gwangju Yeonghwa School, Gyeongbuk Andong Yeongmyeong School, Gyeongnam Jinju Hyegwang School, Jeju Yeongji School, Incheon Cheonghak School, Gyeonggi Ansan Jahae School, Ulsan Meari School, and Sejong Sarang School) serve as the standard's KSL education hubs.
This standard is published as an open standard under the MIT license; all simulator code, specifications, and example code are openly available at GitHub WIA-Official/wia-standards-public/tree/main/sign-language. Our hope is that this standard enables Deaf people who use Korean Sign Language to communicate more freely, and that this standard becomes the foundation for sign-language recognition standards across the Asia-Pacific region.
WIA-Official/wia-standards-public/tree/main/sign-language — Reference implementations of the ELAN template, SignBank export, HamNoSys parser, and SignWriting renderer defined in this chapter are openly published in this repository.