flbench a very unserious benchmark by model Guess the model

Same dare, every model.
You spot who made what.

Each toy here is one prompt given to a pile of AI models — one try, no retries, published exactly as written. Play their takes side by side, then test yourself: can you tell a Fable from a Sonnet by feel alone?

Just for fun. Please don't pick your production model based on a kanban board.

Or take one model at a time

Every toy a single model built, in one place — then flip it against any other model, toy by toy.

One frozen prompt

Every model gets the exact same words — read them on any toy's page and run the dare on your own model if you like.

One try, no edits

Fresh session, no follow-ups, no fixes. What the model wrote in one pass is what you're playing — including the broken bits.

You're the judge

No scores, no rankings, no AI graders. You play the toys and draw your own conclusions. The shelf keeps growing whenever we think of a new dare.