Admin 12 Jun 2026 13:40

 

Sindhi Annotated Corpus

1. Introduction

The Sindhi Annotated Corpus (SAC) is a structured collection of Sindhi language texts that have been manually marked with linguistic information. It provides researchers, developers, and educators with a reliable resource for studying the syntax, morphology, semantics, and pragmatics of Sindhi, a language spoken by around 30 million people in Pakistan and India.

Because Sindhi resources are comparatively scarce, the creation of a highquality annotated corpus represents a significant step toward enabling natural language processing (NLP) applications such as partofspeech tagging, namedentity recognition, machine translation, and sentiment analysis.

2. Corpus Composition

2.1 Text Sources

  • Literary Works: Selections from classic poetry, modern short stories, and novels.
  • News Articles: Articles from leading Sindhi newspapers and online portals, covering politics, economics, culture, and sports.
  • Academic Texts: Abstracts and excerpts from university theses, research papers, and educational material.
  • Social Media: Publicly available posts and comments from platforms such as Twitter, Facebook, and local forums, providing contemporary colloquial language.

2.2 Size and Statistics

As of the latest release (2025), SAC contains approximately 1.2 million tokens, distributed as follows:

  • Literary Works 350,000 tokens
  • News Articles 450,000 tokens
  • Academic Texts 200,000 tokens
  • Social Media 200,000 tokens

The corpus is balanced across genres to support both formal and informal language research.

3. Annotation Layers

The SAC uses a multilayered annotation scheme that follows international standards where possible, while also addressing languagespecific phenomena.

3.1 Tokenisation

Tokens are defined as words, punctuation marks, and symbols. Special attention is given to Sindhi-specific orthographic characters, including the use of the Arabicderived script and the Devanagari variant.

3.2 PartofSpeech (POS) Tagging

Each token receives a POS label from a tagset adapted from the Universal POS Tagset, extended to capture Sindhispecific categories such as postpositions and honorific particles.

3.3 Morphology

Morphological annotation includes:

  • Lemma
  • Root
  • Inflectional features (tense, aspect, mood, gender, number, case)
  • Derivational affixes

3.4 Syntactic Dependencies

Dependency relations are expressed in the CoNLLU format, enabling easy conversion to treebanks. Relations such as subject, object, adverbial, and clausalconnective are annotated.

3.5 NamedEntity Recognition (NER)

Entities are marked using the BIO scheme with categories: PERSON, LOCATION, ORGANIZATION, DATE, TIME, and MISC. An additional category, CULTURAL, captures festival names, traditional clothing, and regional cuisines.

3.6 Sentiment & Pragmatic Tags

For a subset of socialmedia data, sentiment polarity (positive, neutral, negative) and speechact labels (question, request, assertion, exclamation) are provided.

4. Annotation Process

The annotation workflow combines automated preannotation with manual verification:

  1. Preprocessing: Texts are cleaned, normalised, and tokenised using a custom Sindhi tokenizer.
  2. Automatic Tagging: A baseline POS tagger and morphological analyzer generate initial tags.
  3. Manual Review: Trained linguists verify and correct tags using the WebAnno interface.
  4. Quality Assurance: Interannotator agreement is measured (average Cohens =0.87). Discrepancies trigger a secondary review.
  5. Final Export: The annotated data are exported in both CoNLLU and XML formats.

The project follows the LREC guidelines for corpus documentation and sharing.

5. Access and Licensing

The corpus is hosted on the GitHub repository of the Sindhi Language Initiative. It is released under the Creative Commons AttributionNonCommercial 4.0 International (CC BYNC 4.0) license, allowing academic and nonprofit use with appropriate attribution.

Users can download:

  • Full dataset (all layers)
  • Individual layers (e.g., only POS tags)
  • Sample subsets for quick experimentation

Documentation, annotation guidelines, and scripts for data conversion are included in the repository.

6. Applications and Research UseCases

The SAC has already supported a range of projects:

  • POS Tagger Development: A BiLSTMCRF model trained on SAC achieved 94% accuracy on a heldout test set.
  • Machine Translation: SAC serves as a parallel source for SindhiEnglish translation models, improving BLEU scores by 2.5 points over baseline.
  • Sentiment Analysis: Annotated socialmedia data enable sentiment classifiers that reach F1score=0.88 for the negative class.
  • Dialect Studies: By comparing lexical choices across regions, researchers have identified systematic variations between Karachi and Hyderabad dialects.
  • Educational Tools: Languagelearning apps use SACs lemma and POS information to generate exercises.

7. Future Directions

Planned enhancements for the next version of the corpus include:

  • Expanding the token count to 2.5million, with more spokenlanguage recordings.
  • Adding a discourselevel annotation layer for coreference and rhetorical relations.
  • Integrating multimodal data (images with captions) to support visionlanguage research.
  • Developing a webbased query interface for selective download of annotated slices.
  • Collaborating with international bodies to map SAC to the Global Language Resource Repository (GLRR).

Community contributions are encouraged; a contribution guide is available in the repository.

8. References & Further Reading

  • Ali, S., & Khan, R. (2024). Building the First Sindhi Annotated Corpus. Proceedings of LREC 2024, pp. 11231130.
  • Rahman, M. (2023). Morphological Tagging for LowResource Languages: A Sindhi Case Study. Journal of NLP Research, 18(2), 4560.
  • WebAnno A Flexible Annotation Tool. https://webannotation.github.io/
  • Creative Commons Licenses CC BYNC 4.0. https://creativecommons.org/licenses/by-nc/4.0/

Reference Files For Sindhi Annotated Corpus
Screenshoot
File Name
e6c81fad66b4f24ea6e6c37abca41d2457e7.pdf

File Size
1.21 MB

File Type
PDF

File Site
Description
This file is just a reference file for Sindhi Annotated Corpus. Does not guarantee that the specific things you want are included in it.
Direct download (wait 10 seconds)

Sindhi Annotated Corpus and Reference File Download Link


admin
Admin
2026-06-12 13:40:11

Modifications In The Annotated Templates Comparing To 2.2.0 Hotfix and Reference File Down...


admin
Admin
2026-06-06 19:12:15

Annotated Bibliography and Reference File Download Link


admin
Admin
2026-06-08 09:12:15

Annotated Catalogue 2015 and Reference File Download Link


admin
Admin
2026-06-09 08:38:05

Subject Verb Agreement In Sindhi And English Syntax and Reference File Download Link


admin
Admin
2026-06-09 10:16:15