description: Tokenizer used for BERT, a faster version with TFLite support.
Tokenizer used for BERT, a faster version with TFLite support.
Inherits From: TokenizerWithOffsets,
Tokenizer,
SplitterWithOffsets,
Splitter, Detokenizer
text.FastBertTokenizer(
vocab=None,
suffix_indicator='##',
max_bytes_per_word=100,
token_out_type=dtypes.int64,
unknown_token='[UNK]',
no_pretokenization=False,
support_detokenization=False,
fast_wordpiece_model_buffer=None,
lower_case_nfd_strip_accents=False,
fast_bert_normalizer_model_buffer=None
)
This tokenizer applies an end-to-end, text string to wordpiece tokenization. It
is equivalent to BertTokenizer for most common scenarios while running faster
and supporting TFLite. It does not support certain special settings (see the
docs below).
See WordpieceTokenizer for details on the subword tokenization.
For an example of use, see https://www.tensorflow.org/text/guide/bert_preprocessing_guide
detokenize(
token_ids
)
Convert a Tensor or RaggedTensor of wordpiece IDs to string-words.
See
WordpieceTokenizer.detokenize
for details.
Note:
FastBertTokenizer.tokenize/FastBertTokenizer.detokenize
does not round trip losslessly. The result of detokenize will not, in general,
have the same content or offsets as the input to tokenize. This is because the
"basic tokenization" step, that splits the strings into words before applying
the WordpieceTokenizer, includes irreversible steps like lower-casing and
splitting on punctuation. WordpieceTokenizer on the other hand is
reversible.
Note: This method assumes wordpiece IDs are dense on the interval [0, vocab_size).
>>> vocab = ['they', "##'", '##re', 'the', 'great', '##est', '[UNK]']
>>> tokenizer = FastBertTokenizer(vocab=vocab, support_detokenization=True)
>>> tokenizer.detokenize([[4, 5]])
<tf.Tensor: shape=(1,), dtype=string, numpy=array([b'greatest'],
dtype=object)>
| Args | |
|---|---|
| `token_ids` | A `RaggedTensor` or `Tensor` with an int dtype. |
| Returns | |
|---|---|
| A `RaggedTensor` with dtype `string` and the same rank as the input `token_ids`. |
split(
input
)
Alias for
Tokenizer.tokenize.
split_with_offsets(
input
)
Alias for
TokenizerWithOffsets.tokenize_with_offsets.
tokenize(
text_input
)
Tokenizes a tensor of string tokens into subword tokens for BERT.
>>> vocab = ['they', "##'", '##re', 'the', 'great', '##est', '[UNK]']
>>> tokenizer = FastBertTokenizer(vocab=vocab)
>>> text_inputs = tf.constant(['greatest'.encode('utf-8') ])
>>> tokenizer.tokenize(text_inputs)
<tf.RaggedTensor [[4, 5]]>
| Args | |
|---|---|
| `text_input` | input: A `Tensor` or `RaggedTensor` of untokenized UTF-8 strings. |
| Returns | |
|---|---|
| A `RaggedTensor` of tokens where `tokens[i1...iN, j]` is the string contents (or ID in the vocab_lookup_table representing that string) of the `jth` token in `input[i1...iN]` |
tokenize_with_offsets(
text_input
)
Tokenizes a tensor of string tokens into subword tokens for BERT.
>>> vocab = ['they', "##'", '##re', 'the', 'great', '##est', '[UNK]']
>>> tokenizer = FastBertTokenizer(vocab=vocab)
>>> text_inputs = tf.constant(['greatest'.encode('utf-8')])
>>> tokenizer.tokenize_with_offsets(text_inputs)
(<tf.RaggedTensor [[4, 5]]>,
<tf.RaggedTensor [[0, 5]]>,
<tf.RaggedTensor [[5, 8]]>)
| Args | |
|---|---|
| `text_input` | input: A `Tensor` or `RaggedTensor` of untokenized UTF-8 strings. |
| Returns | |
|---|---|
| A tuple of `RaggedTensor`s where the first element is the tokens where `tokens[i1...iN, j]`, the second element is the starting offsets, the third element is the end offset. (Please look at `tokenize` for details on tokens.) |