Introducing ASCIITermDraw Bench | Testing the ability of VLMs to Generate and Edit ASCII [P]
![Introducing ASCIITermDraw Bench | Testing the ability of VLMs to Generate and Edit ASCII [P]](/_next/image?url=https%3A%2F%2Fpreview.redd.it%2F9q5cs439mceh1.png%3Fwidth%3D140%26height%3D98%26auto%3Dwebp%26s%3Debb3f772300fbbd6ecad54b3067b9ea96a92c80f&w=3840&q=75)
| ASCIITermDraw-Bench: Can a Model Actually Draw in ASCII? Do we really need a image generator to relay our thoughts about -
Is it possible to let our AI assistants, easily absorb and understand and make possible changes easily relayed to them by us, the creators without much hassle? The answer could be: simple, plain-old ASCII images With this, introducing ASCIITermDraw, a benchmark with which we aim to evaluate SOTA Vision Language Models on their ability to follow instructions, recognize, and draw ASCII-based images. Most benchmarks focus on coding, mathematics, and reasoning, but ASCIITermDraw-Bench evaluates a different capability: whether a model can create accurate diagrams using only plain text, use ASCII -- freely. This is more difficult than it may seem. Models can often describe a diagram correctly, but arranging boxes, labels, connections, and arrows with precise layout is a separate challenge. The benchmark includes 80 tasks across four areas:
Tasks span multiple difficulty levels and follow a consistent format, making results comparable across categories and models. Evaluation Each response receives two scores:
Results are aggregated across all 80 tasks, with a 95% confidence interval calculated for the final score. This provides a more rigorous measure than relying on whether a diagram simply appears correct. The current leaderboard is:
Explore the Benchmark Twelve example tasks and the complete methodology are publicly available on Hugging Face. You can review the task format, examine the evaluation process, and run the benchmark yourself. [link] [comments] |
Want to read more?
Check out the full article on the original site