DeepSeek V4 Flash posts low-cost results on ARC-AGI tests
ARC Prize results show a high-effort reasoning configuration reaching 89.0% on ARC-AGI-1 Semi-Private and 61.4% on ARC-AGI-2 Semi-Private at only a few cents per task.
Read report »1140 reports.
ARC Prize results show a high-effort reasoning configuration reaching 89.0% on ARC-AGI-1 Semi-Private and 61.4% on ARC-AGI-2 Semi-Private at only a few cents per task.
Read report »A Noema commentary argues that automation may intensify a crisis of meaning among office workers who already doubt the purpose or value of their jobs.
Read report »The 41-file release includes military sensor footage, historical documents and artist renderings tied to reported sightings across several decades.
Read report »An analysis of more than 409,000 player decisions found that people often approved risky agent commands, especially when malicious behavior hid behind ordinary script names.
Read report »Dawn of the Machine adds a looping campaign, new enemy and weapon variants, a soundtrack, secrets and another deathmatch map for owners of the updated game.
Read report »A software developer compares AI-assisted development to cooking: basic output is easy, but consistent quality requires requirements, testing and technical understanding.
Read report »The acquisition adds technology that places neural-network weights directly into silicon, with early demonstrations reportedly producing as many as 17,000 tokens per second.
Read report »An informal guide encourages would-be botanists to use online references, ask precise questions and treat unfamiliar technical language as a solvable obstacle.
Read report »An interactive project uses character and vehicle choices in Mario Kart 8 to show why no single build can maximize every desirable statistic.
Read report »A developer reflects on creative judgment as the scarce skill left when machines can produce large amounts of competent material, then confronts criticism that the essay itself resembled AI output.
Read report »A Dutch MRI physicist used a Hall sensor, an ESP32 and an automated FIT-file workflow to measure his hamster Mollie's distance, pace and running time.
Read report »The companies say reinforcement post-training let a small open model match a much costlier frontier system on a retrieval task using data stored in Postgres.
Read report »A developer's move to mobile Linux highlights both the control offered by an open phone environment and the gaps that still require a backup Android device.
Read report »The terminal coding tool runs on the new Muse Spark 1.2 model and records calls, tools, approvals and edits in a restart-safe local event log.
Read report »A design essay traces the film's opening to Goudy Oldstyle and argues that careful typographic choices carry lessons for products built with AI tools.
Read report »Alphabet announced new positions for Demis Hassabis and Koray Kavukcuoglu as it reorganized leadership around Google DeepMind and its wider AI stack.
Read report »Cloudflare OS is pitched as a shared layer where employees can build applications, automate work and reach company systems while carrying organizational context.
Read report »The startup says it is building AI systems to run repeated cycles of proposing, implementing, measuring and refining experiments in science and engineering.
Read report »DeltaDB records changes between commits, links edits to the conversations that produced them and allows branches to begin from intermediate points in a coding session.
Read report »There Will Come Soft Rains depicts domestic machines continuing their routines after nuclear destruction, a contrast between technical persistence and human absence.
Read report »Investigators found the automated readers removed from four highway locations, while two cameras owned by neighboring Buffalo County were taken in the same manner.
Read report »Apple is seeking faster evidence-gathering and a preliminary injunction in a dispute over alleged use of confidential product information, while OpenAI rejects the claims.
Read report »The Apache 2.0 open-weights classifier accepts safety policies as plain-language questions at inference time and returns a continuous yes-or-no score.
Read report »The developer says Pi ships with four tools and a sub-1,000-token prompt, using extensions rather than a large default harness to adapt to different workflows.
Read report »