Standard for the Presentation of Nucleotide and Amino Acid Sequence Listings Using XML
In the field of biotechnology and pharmaceutical patents, the precise documentation of genetic and protein sequences is critical. The World Intellectual Property Organization (WIPO) established Standard ST.26 to modernize how these sequences are presented. This standard replaces the outdated ST.25, transitioning from a text-based format to an eXtensible Markup Language (XML) format.
The Shift to XML
The transition to XML was driven by the need for greater interoperability, searchability, and data integrity. While the previous text-based standards were human-readable, they were prone to formatting errors and were difficult for computer systems to parse reliably across different global patent offices.
XML allows for a structured, hierarchical data format that ensures that every piece of informationsuch as the organism source, sequence length, and the actual sequence of nucleotides or amino acidsis tagged explicitly. This ensures that the data remains consistent regardless of the software used to view or process it.
Key Objective: The primary goal of the XML standard is to enable the seamless exchange of sequence data between patent offices and public databases, reducing the manual labor required for data entry and minimizing errors in sequence interpretation.
Core Components of the Standard
The XML standard defines a specific schema that all sequence listings must follow. This structure generally includes the following key elements:
- Header Information: Details about the applicant, the date of creation, and the software version used to generate the XML file.
- Sequence Identifiers: Unique IDs for each sequence, allowing for easy cross-referencing within the patent application.
- Feature Tables: Detailed descriptions of specific regions of the sequence, such as promoters, coding regions (CDS), or active sites.
- Sequence Data: The actual string of characters representing the DNA, RNA, or protein sequence, stripped of non-essential formatting.
Advantages of ST.26 over ST.25
The adoption of XML provides several technical advantages over the legacy text formats:
- Validation: XML files can be automatically validated against a schema (XSD). This means a patent office can immediately detect if a listing is missing a required field or contains an illegal character before the application is even processed.
- Searchability: Structured data allows for complex queries. Researchers can search for specific motifs or sequence variations across thousands of patents much faster than with flat-text files.
- Scalability: As biological data grows in size and complexity (such as large genomic sequences), XML handles larger datasets more efficiently than legacy formats.
Example Structure
A simplified representation of how sequence data is structured in an XML-based sequence listing looks like this:
<SequenceListing> <Sequence> <Identifier>SEQ ID NO: 1</Identifier> <Molecule>DNA</Molecule> <Length>120</Length> <SequenceData>ATGC...</SequenceData> </Sequence></SequenceListing>
Compliance and Implementation
For practitioners and scientists, compliance with this standard is mandatory for filings in most major jurisdictions. Failure to provide sequence listings in the correct XML format can lead to formal objections or delays in the patent granting process. Most users utilize WIPO Sequence, a free software tool provided by WIPO to help users generate compliant XML files without needing to write code manually.
The implementation process typically involves extracting sequence data from laboratory notebooks or bioinformatics software, inputting it into a compliant generator, and validating the resulting XML file before submission.
Conclusion
The movement toward an XML-based standard for nucleotide and amino acid sequences represents a significant leap in the digitalization of intellectual property. By unifying the presentation of biological data, WIPO has created a global language for genetic disclosure, fostering transparency and efficiency in the global innovation ecosystem.
Reference Files For Standard For The Presentation Of Nucleotide And Amino Acid Sequence Listings Using XML (eXtensible Markup Language)
File Name
2021_14039.pdf
File Size
0.30 MB
File Type
PDF
File Site
Description
This file is just a reference file for Standard For The Presentation Of Nucleotide And Amino Acid Sequence Listings Using XML (eXtensible Markup Language). Does not guarantee that the specific things you want are included in it.
Direct download (wait 10 seconds)
Standard For The Presentation Of Nucleotide And Amino Acid Sequence Listings Using XML (eX...
Admin
2026-06-11 02:20:11
Uptake Rate Measurement Of Some Amino Acids On Normal And Treated Yeast Cells To Xenobioti...
Admin
2026-06-08 17:00:19
Protein And Amino Acid Digestion Characteristics and Reference File Download Link
Admin
2026-06-08 12:46:06
Label-free Amino Acid Identification For De Novo Protein Sequencing Via TRNA Charging And...
Admin
2026-06-09 15:58:05
General Structure Of An Amino Acid and Reference File Download Link
Admin
2026-06-07 06:20:21
We use cookies to enhance your browsing experience and analyze site traffic. By clicking 'Accept all cookies', you agree to the use of these cookies. You can manage your preferences or learn more in our [Privacy Policy/Cookie Policy.