bench-labs developed **GCTokenizer-v1**, which is a multi-lingual tokenizer Available in four sizes: 32K, 65K, 131K and 262K tokens "S, M, L, XL" It utilizes an encoding scheme which allows it to handle characters in any language around the world
General (multi lingual) Consensus (from multiple model tokenizers consensus) Tokenizer
We included an implementation script too, built like BPE- it can encode arbitrary text, most of the time, efficiently
We're building a dataset to study what humans actually consider AI slop.
SlopFinder shows you a random piece of AI-generated text and gives you one simple control: **how slop is it?** No categories. No complicated forms. Just vote and move on.
Every vote helps build the dataset. ๐งฉ
How does it work? Samples are pulled from existing datasets, shown anonymously, and collected into our annotation pool. After enough votes, they're exported to Hugging Face for everyone to use.
This is an early MVP, so the dataset is small and the system is still evolving.