Ask an AI to count bounces in a video, and it often gets the number wrongاز یک AI بخواهید تعداد برخوردها را در یک ویدیو بشمارد؛ اغلب عدد را اشتباه می‌گوید

Leading AI systems that watch video often miscount simple repeated events, such as a ball hitting a wall. Across 2,190 short test videos, the best system gave the right count only about 21% of the time.سیستم‌های پیشرفته AI که ویدیو را تحلیل می‌کنند، اغلب در شمارش رویدادهای تکراری ساده — مثل برخورد یک توپ به دیوار — دچار اشتباه می‌شوند. در بررسی 2,190 ویدیوی کوتاه آزمایشی، بهترین سیستم فقط در حدود 21% موارد عدد درست را ارائه داد.

ترجمهٔ ماشینی است؛ برای دقت به متن اصلی انگلیسی مراجعه کنید.

Why it matters

Video AI is already being pointed at security cameras, sports clips, factory lines and medical recordings. Many of those jobs need plain counting: how many times did this happen, and when? This study shows that basic bookkeeping breaks down once events come quickly or pile up, even in clean cartoon-like videos with no clutter. It also found something uncomfortable: sometimes a system gave the right total while listing the wrong moments, so a correct answer is not proof it really tracked what happened. These are specific systems tested at one moment in time, and new versions appear constantly, so the exact numbers will change.AI ویدیویی از هم‌اکنون برای تحلیل دوربین‌های امنیتی، کلیپ‌های ورزشی، خطوط تولید کارخانه و تصاویر پزشکی به کار گرفته می‌شود. بسیاری از این کاربردها به شمارش ساده نیاز دارند: این اتفاق چند بار رخ داده و کِی؟ این پژوهش نشان می‌دهد همین حسابداری ابتدایی، وقتی رویدادها سریع اتفاق می‌افتند یا روی هم انباشته می‌شوند، از هم می‌پاشد — حتی در ویدیوهای کارتونی و تمیز بدون شلوغی بصری. یافته‌ی نگران‌کننده‌ی دیگر این بود که گاهی سیستم عدد نهایی درستی می‌داد اما لحظه‌های وقوع رویداد را اشتباه فهرست می‌کرد؛ یعنی درست بودن جواب نهایی، دلیلی بر این نیست که سیستم واقعاً روند وقایع را دنبال کرده باشد. این‌ها سیستم‌های مشخصی هستند که در یک برهه‌ی زمانی خاص آزموده شده‌اند، و چون نسخه‌های جدید مدام عرضه می‌شوند، اعداد دقیق تغییر خواهند کرد.

Who's behind it: Sarvesh Baskar, Zikui Cai and colleagues, University of Maryland. Funded by DARPA and the US National Science Foundation, with private support from Open Philanthropy and Apple.

Summary by the Lemma AI · how we grade

Read the original paper (arxiv.org)