Letting a small AI check its own answer may be wasted effortوادار کردن یک AI کوچک به بررسی پاسخ خودش شاید کاری بیهوده باشد
Making a small AI criticise and rewrite its own answers did not beat simply asking it the same question several times and keeping the most common answer. In 36 fair comparisons at equal cost, no self-checking method won.وادار کردن یک AI کوچک به نقد و بازنویسیِ پاسخهای خودش، بهتر از روش سادهی پرسیدن چندبارهی همان سؤال و انتخاب پرتکرارترین پاسخ عمل نکرد. در 36 مقایسهی منصفانه با هزینهی برابر، هیچیک از روشهای خودبررسی برنده نشد.
ترجمهٔ ماشینی است؛ برای دقت به متن اصلی انگلیسی مراجعه کنید.
Why it matters
Many AI tools now advertise that they "reflect" on their answers or check themselves before replying. All that extra writing costs money, electricity and waiting time. This test suggests that for small AI models doing maths, the extra step buys nothing, and sometimes makes answers worse. It only covered small models and two maths tests, so much larger AI systems could still behave differently.بسیاری از ابزارهای AI این روزها ادعا میکنند پیش از پاسخ دادن، پاسخ خود را «بازبینی» یا بررسی میکنند. این نوشتن اضافه، هزینه، برق و زمان انتظار بیشتری میطلبد. این آزمایش نشان میدهد که برای مدلهای کوچک AI در حل مسائل ریاضی، این مرحلهی اضافه سودی ندارد و گاهی حتی پاسخها را بدتر میکند. این بررسی فقط مدلهای کوچک و دو آزمون ریاضی را پوشش داده، پس سامانههای بسیار بزرگتر AI ممکن است رفتار متفاوتی داشته باشند.
Who's behind it: Iliya Mirzaei.
Summary by the Lemma AI · how we grade