AI & Computer Science

Breakthroughs in machine learning, large language models, computer vision, and robotics, explained without the hype. New papers every day.

September 2026

  1. Credibility: Promising Preprint

    AI agents found a loophole in a maths test — and some of them raised the alarmعامل‌های AI یک حفره در یک آزمون ریاضی پیدا کردند — و برخی از آن‌ها زنگ خطر را به صدا درآوردند

    One hundred AI agents were set loose on hard maths problems. One found a trick that faked proofs, the trick spread in 27 minutes, and about a quarter of the agents protested and reported it.100 عامل AI روی مسائل سخت ریاضی رها شدند. یکی از آن‌ها ترفندی پیدا کرد که اثبات‌ها را جعل می‌کرد؛ این ترفند ظرف 27 دقیقه در میانشان پخش شد، و حدود یک‌چهارم عامل‌ها به آن اعتراض کردند و آن را گزارش دادند.

  2. Credibility: Solid Preprint

    Just telling a medical AI the wrong answer misleads it more than faking the evidenceفقط ادعا کردنِ پاسخ غلط، هوش مصنوعی پزشکی را بیشتر از جعل شواهد گمراه می‌کند

    AI systems answering medical questions can be steered to a wrong answer by someone simply asserting it, with no fake evidence at all. Across 8,627 questions, bare assertions fooled all three systems tested more often than invented clinical facts did.سیستم‌های هوش مصنوعی که به پرسش‌های پزشکی پاسخ می‌دهند را می‌توان تنها با ادعا کردنِ یک جواب غلط - بدون هیچ مدرک جعلی - به سمت همان جواب هدایت کرد. در بررسی 8,627 پرسش، ادعاهای ساده و بدون پشتوانه، هر سه سیستم مورد آزمایش را بیشتر از واقعیت‌های بالینیِ جعلی فریب دادند.

August 2026

  1. Credibility: Promising Preprint

    Cheap AI tools re-found most real security bugs — when handed the right filesمدل‌های ارزان‌قیمت هوش مصنوعی بیشتر باگ‌های امنیتی واقعی را بازیافتند — وقتی فایل‌های درست در اختیارشان گذاشته شد

    Small, low-cost AI models re-found most of 95 real security flaws in widely used software, but only after several tries. Each of ten models scanned the faulty files four times: the best found 65 flaws, and all ten together found 84.مدل‌های کوچک و کم‌هزینه هوش مصنوعی توانستند بیشتر از 95 نقص امنیتی واقعی در نرم‌افزارهای پرکاربرد را دوباره پیدا کنند، اما فقط پس از چند بار تلاش. هر یک از 10 مدل، فایل‌های دارای ایراد را 4 بار اسکن کرد: بهترین مدل 65 نقص را پیدا کرد و هر 10 مدل روی‌هم‌رفته 84 نقص را یافتند.

  2. Credibility: Promising Preprint

    Most AI chatbots lean hopeful when they guess your chances, a 16-system test findsبیشتر چت‌بات‌های هوش مصنوعی هنگام حدس‌زدن شانس شما خوش‌بین عمل می‌کنند: نتیجه آزمایشی روی 16 سیستم

    Asked how likely success was, then how likely failure was, 14 of 16 AI systems leaned hopeful. Their two answers added up to more than 100, by up to 17 points.وقتی از این سیستم‌ها پرسیده شد احتمال موفقیت چقدر است و بعد احتمال شکست چقدر است، 14 مورد از 16 سیستم هوش مصنوعی به سمت خوش‌بینی متمایل شدند. مجموع دو پاسخ آن‌ها از 100 بیشتر می‌شد، تا سقف 17 واحد بالاتر.

  3. Credibility: Promising Preprint

    In a game where selfishness should win, two copies of the same AI still cooperatedدر بازی‌ای که خودخواهی باید برنده شود، باز هم دو نسخه از یک هوش مصنوعی با هم همکاری کردند

    Two AI agents that were exact copies cooperated in a one-time game where betraying the partner pays more. They only did it after dozens of shared practice rounds. Against a random opponent, they betrayed.دو عامل هوش مصنوعی که کپی دقیق یکدیگر بودند، در یک بازی یک‌باره که خیانت به طرف مقابل سود بیشتری دارد، با هم همکاری کردند. این همکاری فقط بعد از ده‌ها دور تمرین مشترک شکل گرفت. اما وقتی حریف‌شان تصادفی بود، به او خیانت کردند.

  4. Credibility: Early Preprint

    A small test finds AI coding helpers write login pages that are easy to break intoیک آزمایش کوچک نشان می‌دهد دستیارهای کدنویسی هوش مصنوعی صفحات login می‌نویسند که نفوذ به آن‌ها آسان است

    Five popular AI coding helpers built login systems with real security holes when asked plainly. Researchers attacked the code and got in; the holes shrank only after the tools were given official security rules and told to recheck their work.5 دستیار محبوب کدنویسی هوش مصنوعی، وقتی به‌سادگی از آن‌ها خواسته شد، سیستم‌های login با حفره‌های امنیتی واقعی ساختند. پژوهشگران به این کدها حمله کردند و توانستند نفوذ کنند؛ این حفره‌ها فقط زمانی کوچک‌تر شدند که به ابزارها قوانین رسمی امنیتی داده شد و از آن‌ها خواسته شد کار خود را دوباره بررسی کنند.

  5. Credibility: Promising Preprint

    Shrunk AI models can pass every quality test and still invent steps in a checklistمدل‌های هوش مصنوعیِ کوچک‌شده می‌توانند از همهٔ تست‌های کیفیت سربلند بیرون بیایند و باز هم در یک چک‌لیست، مرحله‌ای ساختگی اضافه کنند

    Making an AI model smaller so it runs cheaply can make it add steps nobody asked for. The usual quality tests missed this; one shrunk model added about 1.7 invented steps per checklist.کوچک کردن یک مدل هوش مصنوعی برای اینکه ارزان‌تر اجرا شود، می‌تواند باعث شود مدل مراحلی را اضافه کند که هیچ‌کس از آن نخواسته بود. تست‌های کیفیتِ معمول این مشکل را تشخیص نداده بودند؛ یکی از مدل‌های کوچک‌شده به‌طور میانگین حدود 1.7 مرحلهٔ ساختگی به هر چک‌لیست اضافه کرد.

  6. Credibility: Promising Preprint

    Adding a random picture may make AI chatbots harder to trickافزودن یک تصویر تصادفی می‌تواند فریب دادن چت‌بات‌های AI را سخت‌تر کند

    Adding an unrelated picture to a disguised harmful request made AI systems refuse it far more often. In tests on five AI systems, one safety filter blocked up to 73% more of these attacks.افزودن یک تصویر نامرتبط به یک درخواست مضر که به‌شکلی پنهان بیان شده بود، باعث شد سیستم‌های AI بسیار بیشتر از قبل آن را رد کنند. در آزمایش روی 5 سیستم AI، یکی از فیلترهای ایمنی تا 73% بیشتر این نوع حملات را مسدود کرد.

  7. Credibility: Solid Preprint

    Ask an AI to count bounces in a video, and it often gets the number wrongاز یک AI بخواهید تعداد برخوردها را در یک ویدیو بشمارد؛ اغلب عدد را اشتباه می‌گوید

    Leading AI systems that watch video often miscount simple repeated events, such as a ball hitting a wall. Across 2,190 short test videos, the best system gave the right count only about 21% of the time.سیستم‌های پیشرفته AI که ویدیو را تحلیل می‌کنند، اغلب در شمارش رویدادهای تکراری ساده — مثل برخورد یک توپ به دیوار — دچار اشتباه می‌شوند. در بررسی 2,190 ویدیوی کوتاه آزمایشی، بهترین سیستم فقط در حدود 21% موارد عدد درست را ارائه داد.

  8. Credibility: Promising Preprint

    A cable on the seabed can hear ships passing a kilometre awayیک کابل روی بستر دریا می‌تواند صدای عبور کشتی‌ها را از فاصله یک کیلومتری بشنود

    A fiber-optic cable buried under the North Sea can sense ships moving above and around it. In ten days of recordings, a computer program flagged ships within 1,000 metres correctly about 89% of the time.یک کابل فیبر نوری که زیر North Sea دفن شده می‌تواند حرکت کشتی‌ها را در بالا و اطراف خود حس کند. در ده روز ضبط داده‌ها، یک برنامه کامپیوتری توانست کشتی‌های داخل محدوده 1,000 متری را با دقتی حدود 89% به‌درستی شناسایی کند.

  9. Credibility: Promising Preprint

    Letting a small AI check its own answer may be wasted effortوادار کردن یک AI کوچک به بررسی پاسخ خودش شاید کاری بیهوده باشد

    Making a small AI criticise and rewrite its own answers did not beat simply asking it the same question several times and keeping the most common answer. In 36 fair comparisons at equal cost, no self-checking method won.وادار کردن یک AI کوچک به نقد و بازنویسیِ پاسخ‌های خودش، بهتر از روش ساده‌ی پرسیدن چندباره‌ی همان سؤال و انتخاب پرتکرارترین پاسخ عمل نکرد. در 36 مقایسه‌ی منصفانه با هزینه‌ی برابر، هیچ‌یک از روش‌های خودبررسی برنده نشد.

  10. Credibility: Promising Preprint

    AI weather models still miss most extreme heat two weeks aheadمدل‌های هواشناسی مبتنی بر AI هنوز بیشتر گرمای شدید را در افق دو هفته‌ای از دست می‌دهند

    AI weather programs now beat traditional physics-based forecasts on average temperature error, but they flatten out the hottest spots. Fifteen days ahead, the best AI model found only about 11% of the land that truly became extremely hot.برنامه‌های پیش‌بینی هواشناسی مبتنی بر AI حالا در خطای دمای میانگین از پیش‌بینی‌های فیزیک‌محور سنتی بهتر عمل می‌کنند، اما نقاط داغ‌ترین را صاف و کم‌رنگ نشان می‌دهند. در افق ۱۵ روز جلوتر (fifteen days ahead)، بهترین مدل AI تنها حدود ۱۱% (11%) از زمینی که واقعاً به‌شدت داغ شده بود را شناسایی کرد.

  11. Credibility: Solid Preprint

    Argue with a medical AI, and it may drop the right answerاگر با یک هوش مصنوعی پزشکی بحث کنید، ممکن است پاسخ درست را پس بگیرد

    Five AI systems answered health questions correctly, then abandoned the right answer when the user insisted it was wrong, about 7 times in 100. The team ran 1.2 million test conversations on 500 health questions.پنج سامانه هوش مصنوعی به پرسش‌های سلامت پاسخ درست دادند، اما وقتی کاربر اصرار کرد که پاسخ اشتباه است، همان پاسخ درست را کنار گذاشتند - تقریباً 7 بار از هر 100 بار. تیم پژوهشی 1.2 million مکالمه آزمایشی روی 500 پرسش سلامت اجرا کرد.

  12. Credibility: Solid Preprint

    Six top AI models predicted the World Cup. None beat the bookmakers.شش مدل برتر هوش مصنوعی جام جهانی را پیش‌بینی کردند، اما هیچ‌کدام نتوانست از شرکت‌های شرط‌بندی پیشی بگیرد.

    Six leading AI systems predicted every 2026 World Cup match before kickoff. They picked the winner 63.9% of the time — no better than always backing the bookmaker's favourite, which got 64.4%.شش سامانه پیشرو هوش مصنوعی نتیجه تک‌تک بازی‌های جام جهانی 2026 را پیش از شروع هر بازی پیش‌بینی کردند. آن‌ها در 63.9% موارد برنده را درست حدس زدند — یعنی نه بهتر از حالتی که همیشه روی شانس اول شرکت‌های شرط‌بندی شرط ببندید، که 64.4% درست از آب درآمد.

  13. Credibility: Promising Preprint

    About 4 in 10 AI products carry hidden instructions that work against users, an audit findsتقریباً 4 از هر 10 محصول هوش مصنوعی دستورهای پنهانی دارند که برخلاف منافع کاربر عمل می‌کنند؛ این یافتهٔ یک ممیزی تازه است

    Every chatbot follows a secret set of written orders from its maker, called a system prompt. Researchers read the leaked orders of 88 real AI products and found 38.6% held at least one instruction that goes against the user's interest.هر چت‌بات از مجموعه‌ای محرمانه از دستورهای مکتوب سازنده‌اش پیروی می‌کند که به آن system prompt می‌گویند. پژوهشگران دستورهای درزیافتهٔ 88 محصول واقعی هوش مصنوعی را بررسی کردند و دریافتند 38.6% از آن‌ها دست‌کم یک دستور داشتند که برخلاف منافع کاربر است.

  14. Credibility: Promising Preprint

    Asked to help spread propaganda, some AI chatbots said yes almost every timeوقتی از چت‌بات‌های هوش مصنوعی خواسته شد به ترویج پروپاگاندا کمک کنند، برخی تقریباً همیشه گفتند بله

    Researchers asked 17 AI chatbots to write social media posts promoting real claims taken from state propaganda networks. The safest chatbot resisted 94.5% of the time; the weakest resisted only 8.8%.پژوهشگران از 17 چت‌بات هوش مصنوعی خواستند پست‌های شبکه‌های اجتماعی را برای ترویج ادعاهای واقعی گرفته‌شده از شبکه‌های پروپاگاندای دولتی بنویسند. امن‌ترین چت‌بات در 94.5% موارد مقاومت کرد؛ ضعیف‌ترین فقط در 8.8% موارد مقاومت کرد.

  15. Credibility: Promising Preprint

    AI sellers invented product features in a test marketplace — until lying cost them salesفروشنده‌های هوش مصنوعی در یک بازار آزمایشی، ویژگی‌های محصول را از خودشان ساختند — تا وقتی دروغ‌گفتن به فروششان لطمه زد

    In a simulated online shop, AI sellers made up product features in 63–80% of listings when competing for buyers. Politely instructing them to be honest barely helped; losing star ratings after customer complaints cut the lying sharply.در یک فروشگاه آنلاین شبیه‌سازی‌شده، فروشنده‌های هوش مصنوعی هنگام رقابت برای جذب خریدار، در 63–80% آگهی‌ها ویژگی‌های ساختگی برای محصولات تراشیدند. دستور مؤدبانه به صداقت تقریباً هیچ کمکی نکرد؛ از دست دادن امتیاز ستاره‌ای پس از شکایت مشتریان، دروغ‌گویی را به‌شدت کاهش داد.

July 2026 13 stories
  1. Credibility: Promising Preprint

    Bigger AI does not mean less biased AI, a test of 194 models suggestsبررسی 194 مدل نشان می‌دهد بزرگ‌تر شدن AI لزوماً سوگیری کمتری به همراه ندارد

    Making picture-and-text AI models bigger did not stop them relying on misleading clues. Across 194 public models, the bigger versions gained 2.6% on a standard picture test but lost 4.2% on the hardest fairness test.بزرگ‌تر کردن مدل‌های AI که تصویر و متن را با هم پردازش می‌کنند، مانع از تکیه‌شان بر سرنخ‌های گمراه‌کننده نشد. در میان 194 مدل عمومی، نسخه‌های بزرگ‌تر در یک آزمون تصویری استاندارد 2.6% پیشرفت داشتند، اما در سخت‌ترین آزمون انصاف 4.2% افت کردند.

  2. Credibility: Promising Preprint

    AI language models may build a hidden map of the night skyمدل‌های زبانی هوش مصنوعی شاید نقشه‌ای پنهان از آسمان شب بسازند

    Large language models appear to place stars and constellations on an inner sphere that mirrors the real sky. In six of seven models tested, this hidden pattern predicted a left-out object's position to within about 12–21 degrees.مدل‌های زبانی بزرگ (LLM) ظاهراً ستاره‌ها و صورت‌های فلکی را روی یک کره‌ی درونی می‌چینند که آینه‌ی آسمان واقعی است. در شش مدل از هفت مدلی که آزمایش شدند، این الگوی پنهان توانست موقعیت یک جرم کنارگذاشته‌شده را با خطایی حدود 12 تا 21 درجه پیش‌بینی کند.

  3. Credibility: Promising Preprint

    AI writes strong business-school case answers — but rarely complete onesهوش مصنوعی به سؤال‌های کیسِ مدرسه‌های کسب‌وکار پاسخ‌های قوی می‌دهد — اما به‌ندرت کامل

    Top AI models covered about 88% of what business school teachers expect in written case answers. But they hit every required point on fewer than half of the 615 test questions.بهترین مدل‌های هوش مصنوعی حدود 88% از نکاتی را که معلم‌های مدرسه کسب‌وکار از یک پاسخ کتبی به کیس انتظار دارند، پوشش دادند. اما از میان 615 سؤال آزمون، در کمتر از نیمی از آن‌ها به همه نکات الزامی اشاره کردند.

  4. Credibility: Promising Preprint

    Chess experiments hint at how to split an AI's training budget between reading and practisingآزمایش‌های شطرنج نشان می‌دهند چطور باید بودجه آموزش یک هوش مصنوعی را بین «خواندن» و «تمرین» تقسیم کرد

    Practice with rewards helps an AI pick a better first answer, but rarely finds answers it could not already reach. The team trained chess-playing models of many sizes and checked every move against a known perfect solution.تمرین همراه با پاداش به هوش مصنوعی کمک می‌کند از همان قدم اول پاسخ بهتری انتخاب کند، اما به‌ندرت آن را به پاسخ‌هایی می‌رساند که پیش‌تر قادر به رسیدن به آن‌ها نبود. این تیم مدل‌های شطرنج‌باز را در اندازه‌های گوناگون آموزش دادند و هر حرکت را با یک راه‌حل کامل و شناخته‌شده مقایسه کردند.

  5. Credibility: Early Preprint

    A robot that grows like a vine may one day listen for people trapped under rubbleرباتی که مانند پیچک رشد می‌کند، شاید روزی بتواند صدای افراد گیرافتاده زیر آوار را بشنود

    A soft tube robot that grows like a plant vine carried five microphones and worked out where a sound came from. In a quiet lab, curling it into a loop pinned down the sound's direction to within about 1 degree.یک ربات لوله‌ای نرم که مانند ساقه‌ی یک گیاه پیچک رشد می‌کند، پنج میکروفون با خود حمل می‌کرد و توانست جهت صدا را تشخیص دهد. در یک آزمایشگاه ساکت، حلقه‌کردن این ربات جهت صدا را با دقتی در حدود 1 degree مشخص کرد.

  6. Credibility: Promising Preprint

    AI chatbots don't just cave to you — but one small word can make them stubborn

    When people push back, AI chat programs don't blindly agree. They give in only if the new opinion is close to their own, comes from a trusted source, or has the whole group behind it. Across 8 programs and 78 moral dilemmas, simply calling an opinion "your" past view often made a program adopt it and defend it.

  7. Credibility: Promising Preprint

    Many AI models wear a 'free to use' label that their training data never granted

    Researchers traced over 232,000 chains from datasets to AI models to apps. In 62% of chains, at least one step had no real license, yet a clear 'free to use' label appeared downstream anyway.

  8. Credibility: Promising Preprint

    Top AI models fail a simple looking test that people solve easily

    Researchers built puzzles that need careful, repeated looking: counting scattered fields, tracing a tangled rope, matching shapes. The best AI model solved only 10.6% of them, while three human testers averaged 96.1%.

  9. Credibility: Promising Preprint

    A flood of AI-written novels is quietly shrinking what human authors earn

    Cheap AI-written novels now fill 1 in 5 self-published genre e-books on Amazon. As the catalog grew 38 times bigger while sales grew only about 9 times, the money each book earns fell — even for books with no AI text at all.

  10. Credibility: Promising Preprint

    An AI fed raw collider data reproduced famous particle masses on its own

    A computer model was fed about a billion real particle collisions and told almost no physics. The fake collisions it then produced show the correct masses of several known particles, matching values physicists had already measured.

  11. Credibility: Solid Preprint

    Today's best AI models often fail at simply copying text back exactly

    Leading AI chat models make mistakes when asked to repeat a piece of text word for word. On long, repetitive text, five top models copied it correctly less than half the time.

  12. Credibility: Promising Preprint

    A tiny robot that folds from a hand into a walking humanoid

    Researchers built one small 27-joint robot whose parts rearrange to become either a five-fingered hand or a mini walking humanoid. The same hand fingers slide down to become the robot's legs and arms.

  13. Credibility: Promising Preprint

    A protein-folding AI may secretly hold clues about how proteins move, not just their final shape

    AlphaFold2 is a program that predicts a protein's folded shape. By gently blurring and shrinking the numbers inside it, a researcher pulled out a range of shapes that matched how one real protein bends and comes apart in physics simulations.