Chess experiments hint at how to split an AI's training budget between reading and practisingآزمایشهای شطرنج نشان میدهند چطور باید بودجه آموزش یک هوش مصنوعی را بین «خواندن» و «تمرین» تقسیم کرد
Practice with rewards helps an AI pick a better first answer, but rarely finds answers it could not already reach. The team trained chess-playing models of many sizes and checked every move against a known perfect solution.تمرین همراه با پاداش به هوش مصنوعی کمک میکند از همان قدم اول پاسخ بهتری انتخاب کند، اما بهندرت آن را به پاسخهایی میرساند که پیشتر قادر به رسیدن به آنها نبود. این تیم مدلهای شطرنجباز را در اندازههای گوناگون آموزش دادند و هر حرکت را با یک راهحل کامل و شناختهشده مقایسه کردند.
ترجمهٔ ماشینی است؛ برای دقت به متن اصلی انگلیسی مراجعه کنید.
Why it matters
Training a big AI costs a lot of electricity, money and computer time, and companies argue about where to spend it. This study offers a rough rule: let the model read a large amount of data first, then practise with rewards, and give practice a bigger share as the total budget grows. It also shows a limit worth knowing: practice mostly sharpened choices the model already leaned towards, and on hard problems it sometimes strengthened wrong answers. These were small chess models, so full-size language models may behave differently.آموزش یک هوش مصنوعی بزرگ برق و پول و زمان محاسباتی زیادی میطلبد، و شرکتها بر سر نحوه هزینهکردن آن اختلافنظر دارند. این پژوهش یک قاعده تقریبی پیشنهاد میکند: نخست بگذارید مدل حجم زیادی داده بخواند، سپس با پاداش تمرین کند، و هرچه بودجه کلی بزرگتر شود، سهم تمرین را بیشتر کنید. این پژوهش محدودیتی هم نشان میدهد که دانستنش مهم است: تمرین بیشتر گرایشهایی را که مدل از پیش داشت تقویت میکرد، و در مسائل دشوار گاهی حتی پاسخهای نادرست را تقویت میکرد. اینها مدلهای کوچک شطرنج بودند، پس ممکن است مدلهای زبانی در اندازه کامل رفتار متفاوتی داشته باشند.
Who's behind it: Jingyan Shen, Ang Li, Pavel Izmailov and colleagues, New York University, with UCLA, Columbia University, University of Illinois Urbana-Champaign and Modal Labs. No outside funding declared; university computing time was provided.
Summary by the Lemma AI · how we grade