Page History

Excerpt
Splits

...

text blocks into separate tokens on any number of white

...

spaces.

Operates On: Lexical Items with TEXT_BLOCK and possibly other flags as specified below.

Saga_is_recognizer

Recognizer	false

Include Page

	Generic Configuration Parameters
	Generic Configuration Parameters

Example Output

Code Block

language	js
title	Example Configuration
text

V-----------[This is a sentenceA.]-----------V----------------------[  and this is a sentence with leading whitespace.]----------------------V 
^--[This]--V--[is]--V--[a]--V--[sentenceA.]--^--[and]--V--[this]--V--[is]--V--[a]--V--[sentence]--V--[with]--V--[leading]--V--[whitespace.]--^{
 "type":"WhitespaceTokenizer"
}

Output Flags

Lex-Item Flags:

TOKEN - Identifies that the Lex-Items produced by this stage are tokens and not text blocks.
ORIGINAL - Identifies that the Lex-Items produced by this stage are the original, as written, representation of every token (e.g. before normalization).
ALL_LETTERS- All of the characters in the token are characters.
ALL_PUNCTUATION - All of the characters in the token have punctuation.
ALL_DIGITS - All of the characters in the token are digits (0-9)
TOKEN - All tokens produced are tagged as TOKEN
HAS_LETTER - Tokens produced with at least one letter character are tagged as HAS_LETTER
HAS_DIGIT - Tokens produced with at least one digit character are tagged as HAS_DIGIT
HAS_PUNCTUATION - Tokens produced with at least one punctuation character are tagged as HAS_PUNCTUATION. (ALL_PUNCTUATION will not be tagged as HAS_PUNCTUATION)

Vertex Flags:

ALL_WHITESPACE - Identifies that the characters spanned by the vertex are all whitespace characters (spaces, tabs, new-lines, carriage returns, etc.).

Page tree

Versions Compared

Old Version 6

New Version Current

Key

Example Output

Output Flags

Lex-Item Flags:

Vertex Flags: