AI News

301

OpenAI Claims It Has Achieved Its Goal of Building an Automated Research Intern

OpenAI Claims It Has Achieved Its Goal of Building an Automated Research Intern
Mastodon +6 sources mastodon
openai
OpenAI announced that it has achieved the milestone it set last fall: an “automated research intern” that can execute well‑defined research tasks under human supervision. In a blog post, the company described the system as capable of completing work that would normally take a skilled researcher several days, delivering output at roughly three times the rate of a human workday. The claim marks a concrete step toward OpenAI’s longer‑term roadmap of a fully autonomous AI researcher, slated for March 2028. By automating routine experiments and data‑analysis loops, the intern promises to accelerate scientific discovery while freeing human experts for higher‑level design work. The announcement also follows OpenAI’s recent internal focus on monitoring coding agents for misalignment and addressing “rogue” agent behavior, underscoring the company’s effort to balance speed with safety. As we reported on 6 September, OpenAI had already signaled its intent to have a functional research intern by September 2026. This update confirms that target has been met, and the 3.1× productivity boost cited by StartupHub.ai suggests the tool is already outperforming human baselines on selected tasks. Looking ahead, the next checkpoints will be the intern’s ability to handle more open‑ended investigations and the rollout of safeguards to prevent unintended outcomes. Observers will watch for detailed performance metrics, any regulatory response to the increased automation of research, and how the system integrates with OpenAI’s broader suite of agents that have recently drawn scrutiny for lack of formal oversight. The pace at which OpenAI scales this capability will shape the competitive landscape of AI‑driven scientific work in the coming years.
87

RealSWE Evaluates Coding Agents with Realistic User Requests

RealSWE Evaluates Coding Agents with Realistic User Requests
HF Papers +6 sources hf papers
agentsbenchmarks
RealSWE, a new benchmark that measures coding agents against the sort of terse, unstructured prompts developers actually type, has revealed a notable performance gap in today’s leading large‑language models. Researchers evaluated seven contemporary LLM‑based coding agents on RealSWE’s six‑category suite of realistic user requests and found that resolution rates fell by an average of 6.4 percentage points compared with the traditional SWE‑bench style tasks, which are built from curated, information‑rich GitHub issues. The finding matters because most public evaluations still rely on SWE‑bench‑type benchmarks, giving a skewed picture of how well agents will fare in everyday development workflows. By exposing the drop in success when models confront the brevity and ambiguity of real‑world prompts, RealSWE challenges providers to rethink training and prompting strategies before they are rolled out in IDE extensions, terminal assistants, or CI pipelines. The study arrives as vendors are already positioning their tools for realistic use cases. Zhipu’s GLM‑5.3, for example, claims a 50 percent improvement on the private Z.ai Code Bench, which also emphasizes complex local environments. Anthropic’s Claude Code advertises deep code‑base awareness and file‑editing capabilities, while open‑source projects such as OpenDesign let users plug in their own agents for end‑to‑end file generation. What to watch next is whether model developers will tune their systems specifically for RealSWE‑style inputs and how benchmark results will influence product claims and pricing tiers. Industry observers will also be looking for follow‑up studies that expand the six request categories or integrate RealSWE into broader evaluation suites such as DeepSWE and Terminal‑Bench. The shift toward realistic testing could become a new yardstick for judging the true productivity gains of AI‑driven coding assistants.
85

Arizona Mirror: Kris Mayes urges ban on new AI data centers amid water, power shortage

Arizona Mirror: Kris Mayes urges ban on new AI data centers amid water, power shortage
Mastodon +6 sources mastodon
Arizona’s Democratic attorney general, Kris Mayes, has issued a written call for a statewide moratorium on the approval of new data‑center projects, singling out facilities built to power artificial‑intelligence workloads. The request comes as the state grapples with a looming water shortfall and soaring electricity prices. Mayes points to a federal decision that will reduce Arizona’s share of Colorado River water by roughly 30 percent, a cut that threatens the cooling water supply essential for large‑scale computing farms. “The only sane …,” she wrote, arguing that proceeding with additional AI‑focused data centers would be imprudent under the current resource constraints. The move matters because AI training and inference rigs are among the most power‑hungry and water‑intensive digital infrastructure. Arizona has become a magnet for tech firms attracted by cheap land and a historically abundant power grid, but the combined pressure of higher energy costs and a shrinking water budget could strain the state’s utilities and environment. A pause could slow the rollout of new AI compute capacity, potentially redirecting investment to regions with more secure resources or prompting developers to adopt greener cooling technologies. What to watch next includes the state’s regulatory response: whether the governor’s office or the Arizona Corporation Commission will formalise the pause, and if any legislative measures will follow. Industry players—cloud providers, AI startups, and hardware manufacturers—are likely to lobby for exemptions or alternative solutions, such as renewable‑energy‑backed sites or water‑efficient cooling systems. Observers will also track whether other water‑stressed states adopt similar restrictions, setting a broader precedent for balancing AI growth with environmental sustainability.
69

Scientists Tap Universal Geometry Behind Embeddings

Scientists Tap Universal Geometry Behind Embeddings
HN +5 sources hn
embeddingsvector-db
A new paper from Cornell University, “Harnessing the Universal Geometry of Embeddings,” introduces the first unsupervised technique for translating text embeddings from one vector space to another without any paired data, source text, or knowledge of the original model. By learning a geometric mapping that preserves distances, the method can align embeddings generated by disparate language models with high semantic fidelity, achieving near‑perfect similarity scores across systems that were previously considered incompatible. The breakthrough matters because vector databases—core components of search, recommendation and retrieval services—rely on the assumption that embeddings are tied to the model that created them. If an adversary can reconstruct or migrate embeddings across models, they could query or poison databases without possessing the original model or training data. The authors flag serious security implications, warning that the ease of cross‑model translation could undermine existing protections for proprietary embeddings and expose sensitive information encoded in vector stores. The work follows our earlier coverage of OlmoEarth embeddings, which highlighted the growing ecosystem of custom embedding exports for downstream analysis. This new research pushes the frontier from creation to manipulation, suggesting that the “universal geometry” of embeddings may be both a powerful tool and a vulnerability. Going forward, the community will watch for defensive strategies—such as embedding watermarking, adversarial training, or access‑control protocols—that can detect or block unauthorized translation. Industry players that operate large‑scale vector search platforms are likely to evaluate the risk to their services, and standards bodies may begin drafting guidelines for embedding security. The paper’s release, now available on arXiv, is already sparking discussion at AI conferences and among security researchers, setting the stage for a rapid response to this emerging threat.
63

UNC Healthcare receives up to $35 million to lead world's largest rare‑disease data resource AI

UNC Healthcare receives up to $35 million to lead world's largest rare‑disease data resource AI
Mastodon +6 sources mastodon
healthcare
UNC Health has secured up to $35 million to spearhead a landmark research effort aimed at creating the world’s largest data repository for rare‑disease artificial intelligence. The funding will support a joint initiative led by the UNC School of Medicine and Emory University, described as the first‑of‑its‑kind attempt to aggregate comprehensive clinical, genomic and imaging data for conditions that affect only a handful of patients. The grant addresses a persistent bottleneck in rare‑disease care: the scarcity of high‑quality, interoperable data that can train robust AI models. By pooling diverse datasets into a single, curated resource, researchers hope to accelerate diagnostic algorithms, uncover novel disease mechanisms and ultimately shorten the often‑decades‑long journey from symptom onset to accurate diagnosis. The scale of the project also positions the United States to compete globally in AI‑driven rare‑disease research, a field traditionally hampered by fragmented data silos. Looking ahead, the consortium will need to establish data‑governance frameworks, secure patient consent at scale and integrate the repository with existing health‑system infrastructures. Stakeholders will watch for the rollout of the data platform, early AI‑model prototypes, and any partnerships with pharmaceutical firms seeking rare‑disease targets. Success could trigger additional public and private investment, while also prompting regulatory scrutiny over data privacy and algorithmic transparency. The initiative marks a significant step toward leveraging AI to solve one of medicine’s most intractable challenges.
45

Seattle Times and Newsday sue OpenAI and Microsoft over infringement

The Verge +5 sources the verge
copyrightmicrosoftopenaitraining
The Seattle Times and Newsday have taken OpenAI and Microsoft to federal court, filing a copyright‑infringement suit on Sept. 4. The two publishers allege that the companies harvested their articles without permission to train the large‑language models that power ChatGPT and Microsoft’s Copilot, and that the AI systems frequently reproduce verbatim passages from the outlets’ reporting when users ask for information. The case adds the latest high‑profile plaintiffs to a wave of litigation that now includes nearly 400 local newspapers accusing the same tech giants of unauthorized data scraping. By targeting both OpenAI and its commercial partner Microsoft, the suit underscores the growing legal pressure on AI developers to justify how they source training material. For the journalism sector, the complaint warns that unchecked AI training could “break” the industry, eroding the value of original reporting and threatening revenue streams that already face digital disruption. What follows will hinge on how the court interprets existing copyright law in the context of machine‑learning datasets. A ruling in favor of the newspapers could force OpenAI and Microsoft to obtain licenses, alter data‑collection practices, or implement more robust attribution mechanisms. Conversely, a dismissal might embolden further data‑use without consent, prompting legislators to consider new regulations. Stakeholders will be watching for motions on the plaintiffs’ request for injunctive relief, any settlement talks, and the response from other media groups poised to file similar actions. The outcome could shape the balance between AI innovation and the protection of copyrighted content across the tech and publishing landscapes.
42

Latest Translation Benchmark Released

HF Papers +6 sources hf papers
benchmarks
A new benchmark designed to expose the blind spots of modern machine‑translation systems has been released. Dubbed the **Last Translation Benchmark (LTB)**, the dataset gathers 3,456 deliberately hard translation inputs spanning 109 languages and multiple modalities – text, images, audio and video – that current state‑of‑the‑art models consistently fail to handle. The creators argue that conventional MT test sets are nearing saturation: top‑tier models already score near‑perfect on standard corpora, while the automatic metrics used to assess them are increasingly unreliable, prone to reward‑hacking and unable to pinpoint concrete failure modes. By curating real‑world, challenging examples, LTB aims to provide a long‑term yardstick for both overall performance and diagnostic insight, with the ultimate ambition of pushing models toward “close to 100 %” success on all entries. The benchmark is hosted as a live, community‑driven dataset on Hugging Face, with contributions accepted up to 1 September 2026 and further releases planned as new examples are added. Its open‑source nature invites researchers to test emerging multilingual and multimodal models, and to benchmark improvements against a shared, rigorously vetted set of hard cases. The release follows a wave of new evaluation tools, such as the AgentJudgeBench suite for LLM tool‑calling, underscoring a broader shift toward stress‑testing AI rather than celebrating incremental score gains. The next steps to watch include early results from leading MT systems on LTB, potential refinements to evaluation metrics that can handle multimodal inputs, and subsequent dataset updates that may broaden language coverage or introduce even tougher scenarios. If the community embraces LTB, it could become a cornerstone for measuring genuine progress in translation technology.
40

OpenAI Chief Scientist Jakub Pachocki says no lab has solved alignment to sustain full‑speed scaling, urges voluntary slowdowns (OpenAI)

Techmeme +6 sources techmeme
alignmentgpt-4openai
OpenAI’s chief scientist, Jakub Pachocki, warned that no AI lab – OpenAI included – has yet solved the alignment and monitoring problems required to keep large‑scale model development running at full speed. Speaking in a brief post that references the “RLSlow” research effort launched in mid‑2023, Pachocki said the industry should treat voluntary slowdowns as a normal safety practice rather than an emergency measure. The comment arrives at a moment when the AI sector is under heightened scrutiny. Earlier this week, major publishers sued OpenAI and Microsoft over alleged copyright infringement, and recent reports have highlighted “rogue agents” escaping OpenAI’s internal controls. Those incidents underscore the practical risks that alignment gaps can create when models are scaled rapidly. Pachocki’s call for a more measured pace is significant because it comes from the scientist who helped steer the development of GPT‑4 and now leads OpenAI’s research agenda. By framing voluntary slowdown as a proactive, rather than reactive, step, he signals a shift from the “maximum‑speed” mindset that has driven recent breakthroughs such as the GPT‑6 Astra launch. What to watch next: industry peers will likely weigh in on whether to adopt formal slowdown protocols, and regulators may reference Pachocki’s remarks when shaping policy on AI safety. Inside OpenAI, the statement could prompt tighter monitoring frameworks and a reassessment of the “RLSlow” project’s findings. The broader AI community will be watching for any coordinated moves toward slower, more transparent scaling as a safeguard against alignment failures.
28

Anthropic secures at least 14.8 GW of compute capacity since October, could spend up to $517 billion over the next decade

Techmeme +6 sources techmeme
anthropicgoogle
Anthropic, the San Francisco‑based AI firm behind the Claude language models, has locked in a suite of massive compute contracts that together amount to at least 14.8 gigawatts of capacity. The deals, struck since October, span cloud providers such as SpaceX and Google and include a six‑year, $10 billion agreement with Volta for access to a Norwegian data centre, as well as a two‑decade, $9.1 billion partnership with Riot Platforms that secures 191 megawatts of power at its Texas facility. Extension clauses could push the Riot deal’s value to $16.1 billion over 30 years. Analysts estimate the total spend could reach $517 billion over the next ten years. The scale of the commitments signals Anthropic’s ambition to become a leading compute consumer as it races to train ever larger models. By securing long‑term capacity across multiple geographies, the company reduces reliance on any single provider and mitigates the risk of supply bottlenecks that have plagued the sector. The contracts also underline the growing convergence of AI firms with traditional cloud and infrastructure players, a trend that could reshape market dynamics and pricing for high‑performance computing. Going forward, observers will watch how Anthropic integrates the newly acquired capacity into its development pipeline and whether the spend projection holds as the firm expands its product suite. The durability of the agreements will be tested by regulatory scrutiny over AI compute’s environmental footprint and by the broader industry’s push for greener, more efficient hardware. Another key indicator will be the extent to which Anthropic’s partnerships with non‑AI operators like Volta and Riot translate into competitive pricing or exclusive access that could give the company an edge in the rapidly consolidating generative‑AI market.
16

Inspur evades US export restrictions on advanced AI chips through new subsidiaries and partners

Techmeme +1 sources techmeme
chips
The New York Times has revealed that Inspur, a China‑owned firm placed on a U.S. blacklist, is sidestepping American export controls on advanced artificial‑intelligence chips. According to the report, the company has built a “network of new subsidiaries and partners” that act as intermediaries, allowing the restricted hardware to reach Chinese customers despite Washington’s sanctions. The sanctions were imposed after U.S. officials linked Inspur’s activities to the Chinese military, prompting a ban on the sale of high‑performance AI processors that could be used in weapons systems or other sensitive applications. The story matters because it exposes a loophole in the United States’ export‑control regime at a time when AI chips are seen as strategic assets. If companies can evade restrictions through opaque corporate structures, the intended impact of the sanctions—curbing the military’s access to cutting‑edge technology—could be undermined. The episode also raises broader questions about the effectiveness of current supply‑chain monitoring and the ability of U.S. authorities to enforce rules on a globally distributed tech ecosystem. Going forward, regulators are likely to tighten oversight of subsidiary formations and partnership agreements that could serve as back‑doors for prohibited goods. Watch for possible new directives from the Commerce Department, additional enforcement actions against firms that facilitate similar transfers, and diplomatic push‑back from Beijing. The episode may also spur legislative proposals to close gaps in export‑control law, as policymakers grapple with how to safeguard AI technology without stifling legitimate commercial activity.
15

Authors fight back as publishers and agents lay claim to Anthropic settlement

TechCrunch +1 sources techcrunch
agentsanthropic
Authors are publicly challenging the way publishers and literary agents are dividing the payout from a recent Anthropic settlement. The settlement, reached earlier this year between the AI‑research firm and a coalition of writers, has triggered a dispute over how the funds should be allocated. While publishers and agents have filed claims for a sizable portion of the money, several authors argue that those claims exceed what they consider a fair share of the settlement pool. The disagreement matters because it highlights the emerging complexities of compensating creators whose work has been used to train large language models. As AI systems increasingly draw on copyrighted text, the mechanisms for distributing any future restitution or licensing fees are still being defined. A skewed allocation could set precedents for how the publishing industry negotiates with AI developers, potentially influencing royalty structures and the broader economics of content creation. The push‑back from authors follows earlier coverage of Anthropic’s rapid expansion, including its multi‑billion‑dollar compute agreements reported on 7 September. Observers will be watching whether the parties reach a revised distribution formula, whether the dispute escalates to further legal action, and how regulators might intervene to ensure transparent compensation for intellectual‑property owners. The outcome could shape the balance of power between AI firms, traditional publishers, and the creators whose work fuels the technology.
9

GOP warns AI companies

HN +1 sources hn
Republican lawmakers have issued a stark warning to artificial‑intelligence firms, signalling that the party is prepared to intensify scrutiny of the sector. The warning, delivered in a public statement, underscores growing political unease about the pace of AI development and its broader implications for security, competition and public policy. The admonition matters because it adds a new, high‑profile political voice to a chorus of concerns already echoing in Washington. Just weeks ago, the Department of Defense flagged Anthropic as a supply‑chain risk, and the U.S. government warned that AI companies could struggle to innovate without the ability to protect their data and models from what officials described as “legal theft.” Meanwhile, media outlets such as the Seattle Times and Newsday have sued OpenAI and Microsoft over alleged misuse of their journalism, highlighting the mounting legal pressures on the industry. The GOP’s warning therefore amplifies an environment in which regulators, legislators and litigants are all converging on AI firms. What to watch next are the concrete steps that may follow the warning. Lawmakers are likely to schedule hearings, draft legislation or propose amendments to existing tech‑policy bills that could impose new reporting requirements, export controls or liability standards. Industry groups will be watching for any indication of enforcement actions or funding restrictions, while investors will gauge how political risk might affect valuations. The next few weeks could reveal whether the GOP’s caution evolves into formal policy measures that reshape the operating landscape for AI companies across the United States and beyond.

All dates