Versions Compared

Key

  • This line was added.
  • This line was removed.
  • Formatting was changed.

...

Operates On:  Lexical Items with TOKEN

Configuration

Parameters

  • splitChars (string, optional) - List of characters which should be used to split tokens.
    • If not present, then tokens are split on any sequence of punctuation. 
  • dontSplitChars (string, optional) - List of characters which will NOT be used to split tokens.
    • This is typically used to identify exceptions (characters which are not used to split tokens) when splitChars is missing.
    • These characters are included in the produced tokens.
  • splitFlag (string, optional) - The flag to be put on the vertex between the two tokens.
    • If missing, defaults to ALL_PUNCTUATION.

Examples

Code Block
languagejs
titleExample Configuration 1
{
 "type":"CharacterSplitter",
 "dontSplitChars":"."
}

Splits on all punctuation, except periods.

...

Code Block
languagejs
titleExample Configuration 1
{
  "type":"CharacterSplitter",
  "splitChars":"-",
  "splitFlag":"DASH_SPLIT"
}

(splits tokens dashes)

Output Flags

Lex-Item Flags:

  • TOKEN - All tokens produced are tagged as TOKEN

...