XML Schema for User-defined Indexing Procedure with Location
This section describes additional constraints imposed on the XML document returned by the user-defined lexer indexing procedure when the third parameter is TRUE. The XML document returned must be valid according to the following XML schema:
<xsd:schema xmlns:xsd="http://www.w3.org/2001/XMLSchema">
<xsd:element name="tokens">
<xsd:complexType>
<xsd:sequence>
<xsd:choice minOccurs="0" maxOccurs="unbounded">
<xsd:element name="eos" type="EmptyTokenType"/>
<xsd:element name="eop" type="EmptyTokenType"/>
<xsd:element name="num" type="DocServiceTokenType"/>
<xsd:group ref="DocServiceCompositeGroup"/>
</xsd:choice>
</xsd:sequence>
</xsd:complexType>
</xsd:element>
<!--
Enforce constraint that compMem element must be preceeded by word element
or compMem element for document service
-->
<xsd:group name="DocServiceCompositeGroup">
<xsd:sequence>
<xsd:element name="word" type="DocServiceTokenType"/>
<xsd:element name="compMem" type="DocServiceTokenType" minOccurs="0"
maxOccurs="unbounded"/>
</xsd:sequence>
</xsd:group>
<!-- EmptyTokenType defines an empty element without attributes -->
<xsd:complexType name="EmptyTokenType"/>
<!--
DocServiceTokenType defines an element with content and mandatory attributes
-->
<xsd:complexType name="DocServiceTokenType">
<xsd:simpleContent>
<xsd:extension base="xsd:token">
<xsd:attribute name="off" type="OffsetType" use="required"/>
<xsd:attribute name="len" type="xsd:unsignedShort" use="required"/>
</xsd:extension>
</xsd:simpleContent>
</xsd:complexType>
<xsd:simpleType name="OffsetType">
<xsd:restriction base="xsd:unsignedInt">
<xsd:maxInclusive value="2147483647"/>
</xsd:restriction>
</xsd:simpleType>
</xsd:schema>
Some of the constraints imposed by this XML Schema are as follows:
-
The root element is tokens. This is mandatory. It has no attributes.
-
The root element can have zero or more child elements. The child elements can be one of the following elements: eos, eop, num, word, and
compMem. Each of these represent a specific type of token. -
The
compMemelement must be preceded by a word element or acompMemelement. -
The eos and eop elements have no attributes and must be empty elements.
-
The num, word, and
compMemelements have two mandatory attributes:offand len. Oracle Text will normalize the content of these elements as follows: convert whitespace characters to space characters, collapse adjacent space characters to a single space character, remove leading and trailing spaces, perform entity reference replacement, and truncate to 255 bytes. -
The
offattribute value must be an integer between 0 and 2147483647 inclusive. -
The
lenattribute value must be an integer between 0 and 65535 inclusive.
Table 2-33 describes the element types defined in the preceding XML Schema. Table 2-34 describes the attributes defined in the preceding XML Schema.
Table 34 User-defined Lexer Indexing Procedure XML Schema Attributes
| Attribute | Description |
|---|---|
| off | This attribute represents the character offset of the token as it appears in the document being tokenized. The offset is with respect to the character document passed to the user-defined lexer indexing procedure, not the document fetched by the datastore. The document fetched by the datastore may be pre-processed by the filter object or the section group object, or both, before being passed to the user-defined lexer indexing procedure. The offset of the first character in the document being tokenized is 0 (zero). Offset information follows USC-2 codepoint semantics. |
| len | This attribute represents the character length (same semantics as SQL function The length is with respect to the character document passed to the user-defined lexer indexing procedure, not the document fetched by the datastore. The document fetched by the datastore may be pre-processed by the filter object or the section group object before being passed to the user-defined lexer indexing procedure. Length information follows USC-2 codepoint semantics. |
Sum of off attribute value and len attribute value must be less than or equal to the total number of characters in the document being tokenized. This is to ensure that the document offset and characters being referenced are within the document boundary.
Example
Document: User-defined Lexer.
Tokens:
<tokens>
<word off="0" len="4"> USE </word>
<word off="5" len="7"> DEF </word>
<word off="13" len="5"> LEX </word>
<eos/>
</tokens>