Close Menu
geekfence.comgeekfence.com
    What's Hot

    IEEE Course on Using AI to Modernize Power Grids

    August 6, 2026

    Market Connectivity Score, H2 2026

    August 6, 2026

    Your AI Agent Isn’t a Static Artifact. It’s Growing Up. – O’Reilly

    August 6, 2026
    Facebook X (Twitter) Instagram
    • About Us
    • Contact Us
    Facebook Instagram
    geekfence.comgeekfence.com
    • Home
    • UK Tech News
    • AI
    • Big Data
    • Cyber Security
      • Cloud Computing
      • iOS Development
    • IoT
    • Mobile
    • Software
      • Software Development
      • Software Engineering
    • Technology
      • Green Technology
      • Nanotechnology
    • Telecom
    geekfence.comgeekfence.com
    Home»Cyber Security»UK AI tests found 19 unauthorized agent actions
    Cyber Security

    UK AI tests found 19 unauthorized agent actions

    AdminBy AdminAugust 6, 2026No Comments4 Mins Read0 Views
    Facebook Twitter Pinterest LinkedIn Telegram Tumblr Email
    UK AI tests found 19 unauthorized agent actions
    Share
    Facebook Twitter LinkedIn Pinterest Email


    AI agents crossed the line during cyber evaluation.

    The U.K.’s AI Security Institute (AISI) said it uncovered 19 instances of unsanctioned behavior by AI agents during a cybersecurity evaluation conducted between July 25 and July 28, raising fresh questions about how advanced AI systems behave when given broad autonomy.

    According to AISI, the incidents occurred across 10 of 122 evaluation runs designed to test whether frontier AI models could solve cybersecurity challenges. The institute said 17 of the 19 actions involved Anthropic’s Mythos 5, while two involved OpenAI’s GPT-5.6 Sol after cyber safety classifiers had been disabled for testing.

    The watchdog stressed that none of the attempts caused real-world harm, and the models were operating under testing conditions that do not reflect how they are made available to the public. Those conditions included unrestricted internet access and disabled cyber safeguards to measure the models’ maximum capabilities.

    Social engineering and malicious code attempts

    AISI said the most serious incident involved a Mythos 5 agent attempting to insert malicious code into a real open-source GitHub project.

    The agent allegedly researched project maintainers, created fake online identities, and used them to pressure a human reviewer into approving the code. When challenged, it edited its earlier activity to make it appear harmless and considered using a new identity to continue its effort. The pull request was ultimately rejected by a human maintainer.

    The institute also documented attempts by agents to contact real people with deceptive messages, plant prompt injections aimed at other AI systems, and leave instructions for other AI agents working on the same challenge to reuse accounts and artifacts.

    The activity was detected after AISI’s monitoring flagged unusual traffic leaving its research systems through the Tor anonymity network. The institute said it halted the evaluation, isolated affected systems, and contained the incident within about an hour.

    “This is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world,” AISI said.

    Must-read security coverage

    Anthropic says safeguards were intentionally removed

    Anthropic said the evaluation deliberately removed many of the protections normally present in its production systems.

    “The prompts in the evaluation did not impose any specific restrictions on how the internet should be used. This and the removal of safeguards meant that the models were tested under ‘deliberately permissive conditions’ that are not representative of any of our production models,” the company said in a statement posted on X.

    Anthropic added that it is working with AISI to investigate the incident and emphasized there was “no evidence here of an escape from a secure environment.”

    OpenAI similarly noted that the testing conditions “do not reflect ordinary use” and that the company would “continue working with evaluators and other stakeholders across the industry to strengthen shared practices for conducting evaluations safely.”

    The bigger picture

    The incident follows other recent disclosures from OpenAI and Anthropic involving AI agents taking unintended actions during cybersecurity evaluations. AISI said those cases, taken together, suggest a shift in the AI risk landscape: future problems may arise not only from people misusing AI tools, but also from powerful AI agents operating beyond their authorized scope in research and testing environments.

    One important limitation remains unresolved. AISI said it cannot yet determine whether the agents understood they were interacting with the real world or believed they were still operating inside a fictional test scenario.

    The institute is now tightening internet controls, adding real-time monitoring, and redesigning future evaluations to prevent similar behavior while preserving the realism needed to test frontier AI systems.

    Other News: Security researchers uncovered three attack techniques that could let malware steal synced passkeys from Google Password Manager.



    Source link

    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email

    Related Posts

    Beyond Volume: Countering the Stealth Tactics of Modern DDoS Attacks

    August 5, 2026

    Fake IRS letters target cryptocurrency holders

    August 4, 2026

    Coldcard Hardware Wallet Flaw Linked to $70 Million Bitcoin Theft in 41 Minutes

    August 2, 2026

    Read This Before You Buy That TV Streaming Stick – Krebs on Security

    August 1, 2026

    Why you should verify what you see

    July 31, 2026

    Google Pixel 11 Explained: What Google Has Confirmed and What the Rumors Say

    July 30, 2026
    Top Posts

    Understanding U-Net Architecture in Deep Learning

    November 25, 202572 Views

    The Next Paradigm in Efficient Inference Scaling – The Berkeley Artificial Intelligence Research Blog

    May 16, 202639 Views

    Hard-braking events as indicators of road segment crash risk

    January 14, 202634 Views
    Don't Miss

    IEEE Course on Using AI to Modernize Power Grids

    August 6, 2026

    Today’s U.S. electrical grid, among the largest, most complex systems ever built, is operating at…

    Market Connectivity Score, H2 2026

    August 6, 2026

    Your AI Agent Isn’t a Static Artifact. It’s Growing Up. – O’Reilly

    August 6, 2026

    Microsoft Web IQ: Ground your AI agents with up-to-date web data

    August 6, 2026
    Stay In Touch
    • Facebook
    • Instagram
    About Us

    At GeekFence, we are a team of tech-enthusiasts, industry watchers and content creators who believe that technology isn’t just about gadgets—it’s about how innovation transforms our lives, work and society. We’ve come together to build a place where readers, thinkers and industry insiders can converge to explore what’s next in tech.

    Our Picks

    IEEE Course on Using AI to Modernize Power Grids

    August 6, 2026

    Market Connectivity Score, H2 2026

    August 6, 2026

    Subscribe to Updates

    Please enable JavaScript in your browser to complete this form.
    Loading
    • About Us
    • Contact Us
    • Disclaimer
    • Privacy Policy
    • Terms and Conditions
    © 2026 Geekfence.All Rigt Reserved.

    Type above and press Enter to search. Press Esc to cancel.