LLMs Struggle With Multi-Step Deduction in Clue | dailyai.report
23 stories from today
Research
163d ago
LLMs Struggle With Multi-Step Deduction in Clue
Researchers built a text‑based Clue variant to test LLMs’ deductive reasoning. Six agents, including GPT-4o-mini and Gemini-2.5-Flash played 18 simulated games, winning only four. Fine‑tuning on logic puzzles did not boost accuracy and sometimes increased reasoning volume without precision. The study highlights persistent reasoning gaps in current models.
The Signal
Future work will explore larger datasets and more complex game rules.