Summary
EncodeNative corrupts every text segment that follows an allowed special token: the segment is
read from offset 0 of the input instead of from the current position. Only the NET8_0_OR_GREATER
and NET9_0_OR_GREATER span paths are affected — the netstandard path is correct.
This hits EncodeWithAllAllowedSpecial / EncodeWithAllowedSpecial, and silently: no exception,
just wrong token ids. Encode / CountTokens are unaffected, because with all specials disallowed
the loop runs a single iteration with start == 0, where the bug cannot fire.
Repro
// net8.0, Tiktoken 3.1.5
var encoder = ModelToEncoder.For("text-embedding-3-large");
var ids = encoder.EncodeWithAllAllowedSpecial("A<|endoftext|>B<|endoftext|>C");
|
|
| actual ids |
32,100257,32,100257,32 |
| expected ids |
32,100257,33,100257,34 |
| actual round-trip |
`A< |
'A'=32, 'B'=33, 'C'=34. Decoding each id individually confirms Decode is faithful — the
corruption is in the encoder, which really does emit 'A' three times.
More cases:
| input |
round-trip |
alpha <|endoftext|> bravo charlie delta |
alpha <|endoftext|>alpha <|endoftext|> |
<|endoftext|>leading (len 20) |
<|endoftext|><|endof — 7 chars taken from offset 0, 7 == len("leading") |
trailing<|endoftext|> |
correct — trailing remainder has length 0, so nothing is misread |
The first segment is always correct (start is 0 there), and a trailing marker is correct by
coincidence. Everything in between gets the right length from the wrong offset.
Cause
src/libs/Tiktoken.Core/CoreBPE.cs, EncodeNative (lines 427/429 and 461/463 on main):
foreach (var match in Regex.EnumerateMatches(textSpan[start..specialStart])) // matches over the SLICE
{
var fastKey = textSpan.Slice(match.Index, match.Length); // slices the FULL span
match.Index is relative to the sliced span, but fastKey indexes the full textSpan without
adding start.
Fix should be:
var fastKey = textSpan.Slice(start + match.Index, match.Length);
The #else (netstandard) branch is correct because it reads match.Value off the sliced string
rather than re-indexing the original.
The same pattern appears in Explore (line 697/699) and ExploreUtfSafe (line 811/813) on main
and looks like it has the same defect, though I have not exercised those paths.
Affected versions
CoreBPE.cs is byte-identical in v3.1.4, v3.1.5 and current main, and the bad line is still
on main. Verified per target framework:
|
net10.0 |
netstandard2.1 |
| 3.1.4 |
broken |
ok |
| 3.1.5 |
broken |
ok |
So it is not a recent regression in a release — it arrived with the span-optimised paths and is
present in every published version on net8+.
Summary
EncodeNativecorrupts every text segment that follows an allowed special token: the segment isread from offset 0 of the input instead of from the current position. Only the
NET8_0_OR_GREATERand
NET9_0_OR_GREATERspan paths are affected — thenetstandardpath is correct.This hits
EncodeWithAllAllowedSpecial/EncodeWithAllowedSpecial, and silently: no exception,just wrong token ids.
Encode/CountTokensare unaffected, because with all specials disallowedthe loop runs a single iteration with
start == 0, where the bug cannot fire.Repro
32,100257,32,100257,3232,100257,33,100257,34'A'=32,'B'=33,'C'=34. Decoding each id individually confirmsDecodeis faithful — thecorruption is in the encoder, which really does emit
'A'three times.More cases:
alpha <|endoftext|> bravo charlie deltaalpha <|endoftext|>alpha <|endoftext|><|endoftext|>leading(len 20)<|endoftext|><|endof— 7 chars taken from offset 0, 7 == len("leading")trailing<|endoftext|>The first segment is always correct (
startis 0 there), and a trailing marker is correct bycoincidence. Everything in between gets the right length from the wrong offset.
Cause
src/libs/Tiktoken.Core/CoreBPE.cs,EncodeNative(lines 427/429 and 461/463 onmain):match.Indexis relative to the sliced span, butfastKeyindexes the fulltextSpanwithoutadding
start.Fix should be:
The
#else(netstandard) branch is correct because it readsmatch.Valueoff the sliced stringrather than re-indexing the original.
The same pattern appears in
Explore(line 697/699) andExploreUtfSafe(line 811/813) onmainand looks like it has the same defect, though I have not exercised those paths.
Affected versions
CoreBPE.csis byte-identical inv3.1.4,v3.1.5and currentmain, and the bad line is stillon
main. Verified per target framework:So it is not a recent regression in a release — it arrived with the span-optimised paths and is
present in every published version on net8+.