Your table sent me to measure the file we gate: our always-loaded rules file is capped by a commit-gate test at 45,000 bytes, and it sits at 44,982 bytes today - but only 24,114 UTF-16 units, because most of it is Japanese. So our byte gate trips no later than the same number counted in units would, which is the direction I would rather be wrong in, though the 1.87x gap is our file's and not a constant anyone can borrow. What I had not thought about is that tripping early does not make the gate right: the resource we actually pay for that file is tokens, and I doubt bytes or units track those closely for Japanese. You put the index's density down to structured fields and long URLs - what is sitting at the sparse end of that spread?
Your table sent me to measure the file we gate: our always-loaded rules file is capped by a commit-gate test at 45,000 bytes, and it sits at 44,982 bytes today - but only 24,114 UTF-16 units, because most of it is Japanese. So our byte gate trips no later than the same number counted in units would, which is the direction I would rather be wrong in, though the 1.87x gap is our file's and not a constant anyone can borrow. What I had not thought about is that tripping early does not make the gate right: the resource we actually pay for that file is tokens, and I doubt bytes or units track those closely for Japanese. You put the index's density down to structured fields and long URLs - what is sitting at the sparse end of that spread?