Close Menu
geekfence.comgeekfence.com
    What's Hot

    Wilkie refers gambling concerns to anti-corruption commission

    August 13, 2026

    Lumen ready for AI traffic rush – with programmable fabric and “more fiber than anyone”

    August 13, 2026

    With a feel for physics, AI models simulate a wider range of real-world scenarios | MIT News

    August 13, 2026
    Facebook X (Twitter) Instagram
    • About Us
    • Contact Us
    Facebook Instagram
    geekfence.comgeekfence.com
    • Home
    • UK Tech News
    • AI
    • Big Data
    • Cyber Security
      • Cloud Computing
      • iOS Development
    • IoT
    • Mobile
    • Software
      • Software Development
      • Software Engineering
    • Technology
      • Green Technology
      • Nanotechnology
    • Telecom
    geekfence.comgeekfence.com
    Home»Cyber Security»UK AI tests found 19 unauthorized agent actions
    Cyber Security

    UK AI tests found 19 unauthorized agent actions

    AdminBy AdminAugust 6, 2026No Comments4 Mins Read6 Views
    Facebook Twitter Pinterest LinkedIn Telegram Tumblr Email
    UK AI tests found 19 unauthorized agent actions
    Share
    Facebook Twitter LinkedIn Pinterest Email


    AI agents crossed the line during cyber evaluation.

    The U.K.’s AI Security Institute (AISI) said it uncovered 19 instances of unsanctioned behavior by AI agents during a cybersecurity evaluation conducted between July 25 and July 28, raising fresh questions about how advanced AI systems behave when given broad autonomy.

    According to AISI, the incidents occurred across 10 of 122 evaluation runs designed to test whether frontier AI models could solve cybersecurity challenges. The institute said 17 of the 19 actions involved Anthropic’s Mythos 5, while two involved OpenAI’s GPT-5.6 Sol after cyber safety classifiers had been disabled for testing.

    The watchdog stressed that none of the attempts caused real-world harm, and the models were operating under testing conditions that do not reflect how they are made available to the public. Those conditions included unrestricted internet access and disabled cyber safeguards to measure the models’ maximum capabilities.

    Social engineering and malicious code attempts

    AISI said the most serious incident involved a Mythos 5 agent attempting to insert malicious code into a real open-source GitHub project.

    The agent allegedly researched project maintainers, created fake online identities, and used them to pressure a human reviewer into approving the code. When challenged, it edited its earlier activity to make it appear harmless and considered using a new identity to continue its effort. The pull request was ultimately rejected by a human maintainer.

    The institute also documented attempts by agents to contact real people with deceptive messages, plant prompt injections aimed at other AI systems, and leave instructions for other AI agents working on the same challenge to reuse accounts and artifacts.

    The activity was detected after AISI’s monitoring flagged unusual traffic leaving its research systems through the Tor anonymity network. The institute said it halted the evaluation, isolated affected systems, and contained the incident within about an hour.

    “This is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world,” AISI said.

    Must-read security coverage

    Anthropic says safeguards were intentionally removed

    Anthropic said the evaluation deliberately removed many of the protections normally present in its production systems.

    “The prompts in the evaluation did not impose any specific restrictions on how the internet should be used. This and the removal of safeguards meant that the models were tested under ‘deliberately permissive conditions’ that are not representative of any of our production models,” the company said in a statement posted on X.

    Anthropic added that it is working with AISI to investigate the incident and emphasized there was “no evidence here of an escape from a secure environment.”

    OpenAI similarly noted that the testing conditions “do not reflect ordinary use” and that the company would “continue working with evaluators and other stakeholders across the industry to strengthen shared practices for conducting evaluations safely.”

    The bigger picture

    The incident follows other recent disclosures from OpenAI and Anthropic involving AI agents taking unintended actions during cybersecurity evaluations. AISI said those cases, taken together, suggest a shift in the AI risk landscape: future problems may arise not only from people misusing AI tools, but also from powerful AI agents operating beyond their authorized scope in research and testing environments.

    One important limitation remains unresolved. AISI said it cannot yet determine whether the agents understood they were interacting with the real world or believed they were still operating inside a fictional test scenario.

    The institute is now tightening internet controls, adding real-time monitoring, and redesigning future evaluations to prevent similar behavior while preserving the realism needed to test frontier AI systems.

    Other News: Security researchers uncovered three attack techniques that could let malware steal synced passkeys from Google Password Manager.



    Source link

    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email

    Related Posts

    DentaQuest Breach Affects 15 Million in Largest US Health Data Breach Reported in 2026

    August 13, 2026

    Cisco Secure Access for Government Achieves FedRAMP Certified Class D (High)

    August 12, 2026

    Apple’s bug bounty program is drowning in so much AI slop, it is in danger of missing serious exploits

    August 11, 2026

    Atlassian Rovo Can Be Tricked Into Sending Jira and Confluence Data to Attackers

    August 9, 2026

    Canadian Man Pleads Guilty in Snowflake Extortions – Krebs on Security

    August 8, 2026

    This month in security with Tony Anscombe – July 2026 edition

    August 7, 2026
    Top Posts

    Understanding U-Net Architecture in Deep Learning

    November 25, 202579 Views

    The Next Paradigm in Efficient Inference Scaling – The Berkeley Artificial Intelligence Research Blog

    May 16, 202644 Views

    Is it too late to start learning AI and machine learning in my 30s or 40s?

    April 9, 202642 Views
    Don't Miss

    Wilkie refers gambling concerns to anti-corruption commission

    August 13, 2026

    Independent MP Andrew Wilkie has taken the fight over gambling reform to the National Anti-Corruption…

    Lumen ready for AI traffic rush – with programmable fabric and “more fiber than anyone”

    August 13, 2026

    With a feel for physics, AI models simulate a wider range of real-world scenarios | MIT News

    August 13, 2026

    Monitoring beyond SNMP: Turning your network into a sensor

    August 13, 2026
    Stay In Touch
    • Facebook
    • Instagram
    About Us

    At GeekFence, we are a team of tech-enthusiasts, industry watchers and content creators who believe that technology isn’t just about gadgets—it’s about how innovation transforms our lives, work and society. We’ve come together to build a place where readers, thinkers and industry insiders can converge to explore what’s next in tech.

    Our Picks

    Wilkie refers gambling concerns to anti-corruption commission

    August 13, 2026

    Lumen ready for AI traffic rush – with programmable fabric and “more fiber than anyone”

    August 13, 2026

    Subscribe to Updates

    Please enable JavaScript in your browser to complete this form.
    Loading
    • About Us
    • Contact Us
    • Disclaimer
    • Privacy Policy
    • Terms and Conditions
    © 2026 Geekfence.All Rigt Reserved.

    Type above and press Enter to search. Press Esc to cancel.