Chess experiments hint at how to split an AI's training budget between reading and practisingآزمایش‌های شطرنج نشان می‌دهند چطور باید بودجه آموزش یک هوش مصنوعی را بین «خواندن» و «تمرین» تقسیم کرد

Practice with rewards helps an AI pick a better first answer, but rarely finds answers it could not already reach. The team trained chess-playing models of many sizes and checked every move against a known perfect solution.تمرین همراه با پاداش به هوش مصنوعی کمک می‌کند از همان قدم اول پاسخ بهتری انتخاب کند، اما به‌ندرت آن را به پاسخ‌هایی می‌رساند که پیش‌تر قادر به رسیدن به آن‌ها نبود. این تیم مدل‌های شطرنج‌باز را در اندازه‌های گوناگون آموزش دادند و هر حرکت را با یک راه‌حل کامل و شناخته‌شده مقایسه کردند.

ترجمهٔ ماشینی است؛ برای دقت به متن اصلی انگلیسی مراجعه کنید.

Why it matters

Training a big AI costs a lot of electricity, money and computer time, and companies argue about where to spend it. This study offers a rough rule: let the model read a large amount of data first, then practise with rewards, and give practice a bigger share as the total budget grows. It also shows a limit worth knowing: practice mostly sharpened choices the model already leaned towards, and on hard problems it sometimes strengthened wrong answers. These were small chess models, so full-size language models may behave differently.آموزش یک هوش مصنوعی بزرگ برق و پول و زمان محاسباتی زیادی می‌طلبد، و شرکت‌ها بر سر نحوه هزینه‌کردن آن اختلاف‌نظر دارند. این پژوهش یک قاعده تقریبی پیشنهاد می‌کند: نخست بگذارید مدل حجم زیادی داده بخواند، سپس با پاداش تمرین کند، و هرچه بودجه کلی بزرگ‌تر شود، سهم تمرین را بیشتر کنید. این پژوهش محدودیتی هم نشان می‌دهد که دانستنش مهم است: تمرین بیشتر گرایش‌هایی را که مدل از پیش داشت تقویت می‌کرد، و در مسائل دشوار گاهی حتی پاسخ‌های نادرست را تقویت می‌کرد. این‌ها مدل‌های کوچک شطرنج بودند، پس ممکن است مدل‌های زبانی در اندازه کامل رفتار متفاوتی داشته باشند.

Who's behind it: Jingyan Shen, Ang Li, Pavel Izmailov and colleagues, New York University, with UCLA, Columbia University, University of Illinois Urbana-Champaign and Modal Labs. No outside funding declared; university computing time was provided.

Summary by the Lemma AI · how we grade

Read the original paper (arxiv.org)