It's funny that there are often people who say "this test is stupid because it's a problem with the tokenizer" or "OpenAI intransparently sends your queries to smaller models" or "it would make sense with version numbers", and yeah. Sure. But Sam promised an LLM that's so intelligent it's scary, and what people see instead is a model that gets the answer wrong because it's not even smart enough to understand that you're talking about decimal numbers instead of versions.
Poisoning alt text? This isn't going far enough. I propose we poison everything we write by inserting random words and fucking up the grammar. Will it be harder to read everything? Sure, but who cares, I want to mess up companies data scraping efforts.
/j