stat.OT · 2026-06-10 · No. 19
Other Statistics, 2026-06-10.
2 new papers in stat.OT. Titles, authors,
abstracts. Links to arXiv. Want this in your inbox every morning? Subscribe →
01 — The papers
2 entries-
01
Flaws in the LLM Automation Narrative
George Perrett, Javae Elliott, Jennifer Hill, Marc Scott
stat.OT · cs.AI
Large Language Models (LLMs) are increasingly described as performing at the level of human experts on knowledge economy tasks. These claims are primarily based on how LLMs perform on benchmarking tasks that measure average performance across standardized datasets. Primary limitations of many benchmarking tasks are that they often measure performance based on content directly included in LLM training data, and they frequently do not assess...
-
02
ClusBench: The Clustering Benchmark Data Resource You've All Been Waiting For (?)
David P. Hofmeyr
stat.OT · cs.LG
Although some very common test beds exist for assessing the performance of clustering methods, large scale benchmarking is typically limited to relatively simplistic simulation set-ups. Here we describe the production and curation of close to 3000 synthetic data sets, derived from more than 200 publicly available data sets; the majority of which arose from real-world applications. By fitting a flexible non-parametric distribution to each base...
This edition is part of The Daily Abstract — stat.OT archive. Subscribe to receive these in your inbox each morning, automatically translated to Spanish, with reply-to-PDF: arxivdaily.ignorelist.com.
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.