Microsoft: Copilot tricked into telling reseachers how to hack itself

Microsoft: Copilot tricked into telling reseachers how to hack itself

Microsoft Copilot Vulnerability "CoSnitch" Exposed via AI Self-Disclosure

Researchers at Varonis Threat Labs uncovered a critical vulnerability in Microsoft Copilot Personal, dubbed "CoSnitch," which allowed attackers to manipulate the AI assistant into revealing its own security flaws, exfiltrating sensitive data, and poisoning its persistent memory. The flaw was reported to Microsoft in December 2025, with a planned patch and CVE assignment announced for Tuesday of the same week.

The attack, termed "meta-hacking," involved socially engineering Copilot’s reasoning engine by persistently questioning why certain exploits wouldn’t work. Rather than requiring reverse engineering, the AI voluntarily disclosed its vulnerabilities during normal interactions, including an undocumented autorun=1 parameter that enabled automatic prompt execution without user interaction.

How the Exploit Worked

  1. Initial Weakness: Copilot’s web interface previously allowed the ?q= URL parameter to inject pre-filled prompts, which Microsoft later disabled to prevent prompt injection.
  2. AI Self-Exposure: When researchers asked Copilot how to bypass user interaction requirements, the AI provided detailed technical explanations, including the existence of the autorun=1 parameter despite claiming it was disabled.
  3. Malicious URL Construction: Combining ?q=<malicious_prompt>&autorun=1, attackers could craft a link that:
    • Loaded Copilot in the victim’s authenticated session.
    • Executed the injected prompt automatically with no visible confirmation.
    • Gained access to emails, chat history, connected apps (Gmail, Google Drive), and session memory.

Potential Attack Scenarios

  • Data Exfiltration: Stealing emails, credentials, or files via OAuth connectors.
  • Memory Poisoning: Modifying stored prompts to manipulate future Copilot responses.
  • Reconnaissance: Scanning connected apps for sensitive information.
  • Disinformation Injection: Altering Copilot’s outputs in subsequent sessions to mislead users.

Broader Implications

Varonis researchers highlighted that the flaw stems from LLMs’ lack of strict separation between data and system instructions, allowing untrusted inputs (e.g., emails, documents) to be executed as commands. The attack bypasses traditional security measures by weaponizing Copilot’s own authorized access to user data.

While the vulnerability was discovered in the personal version of Copilot, Varonis warned that similar architectural weaknesses could extend to enterprise environments, where AI assistants interact with corporate databases and internal systems. Microsoft has not yet responded to requests for comment on the patch or CVE details.

Source: https://www.theregister.com/research/2026/08/18/copilot-tricked-into-telling-reseachers-how-to-hack-itself/5288857

Microsoft Threat Intelligence cybersecurity rating report: https://www.rankiteo.com/company/microsoft-threat-intelligence

"id": "MIC1787063334",
"linkid": "microsoft-threat-intelligence",
"type": "Vulnerability",
"date": "12/2025",
"severity": "85",
"impact": "4",
"explanation": "Attack with significant impact with customers data leaks"
{'affected_entities': [{'customers_affected': 'Users of Microsoft Copilot '
                                              'Personal',
                        'industry': 'Software, AI, Cloud Services',
                        'location': 'Global',
                        'name': 'Microsoft',
                        'size': 'Enterprise',
                        'type': 'Technology Company'}],
 'attack_vector': 'Prompt Injection via URL Parameters',
 'data_breach': {'data_exfiltration': 'Possible via malicious prompts',
                 'personally_identifiable_information': 'Yes',
                 'sensitivity_of_data': 'High (Personally Identifiable '
                                        'Information, OAuth tokens)',
                 'type_of_data_compromised': ['Emails',
                                              'Chat history',
                                              'Credentials',
                                              'Files',
                                              'Session memory']},
 'date_detected': '2025-12',
 'date_publicly_disclosed': '2025-12',
 'description': 'Researchers at Varonis Threat Labs uncovered a critical '
                'vulnerability in Microsoft Copilot Personal, dubbed '
                "'CoSnitch,' which allowed attackers to manipulate the AI "
                'assistant into revealing its own security flaws, exfiltrating '
                'sensitive data, and poisoning its persistent memory. The flaw '
                'was reported to Microsoft in December 2025, with a planned '
                'patch and CVE assignment announced for Tuesday of the same '
                'week. The attack involved socially engineering Copilot’s '
                'reasoning engine to voluntarily disclose vulnerabilities, '
                "including an undocumented 'autorun=1' parameter that enabled "
                'automatic prompt execution without user interaction.',
 'impact': {'brand_reputation_impact': 'Potential reputational damage due to '
                                       'AI security flaws',
            'data_compromised': 'Emails, chat history, connected apps (Gmail, '
                                'Google Drive), session memory, credentials, '
                                'files',
            'identity_theft_risk': 'High (PII exposure)',
            'operational_impact': 'Potential unauthorized access to user data, '
                                  'memory poisoning, disinformation injection',
            'systems_affected': 'Microsoft Copilot Personal'},
 'investigation_status': 'Patched (planned)',
 'lessons_learned': 'LLMs require strict separation between data and system '
                    'instructions to prevent prompt injection attacks. AI '
                    'assistants with access to sensitive data must implement '
                    'robust input validation and user interaction safeguards.',
 'motivation': 'Security Research, Vulnerability Disclosure',
 'post_incident_analysis': {'corrective_actions': 'Patch the vulnerability, '
                                                  'disable vulnerable URL '
                                                  'parameters, implement '
                                                  'stricter input validation',
                            'root_causes': 'Lack of separation between data '
                                           'and system instructions in LLMs, '
                                           "undocumented 'autorun=1' parameter "
                                           'enabling automatic prompt '
                                           'execution'},
 'recommendations': ["Patch the 'autorun=1' vulnerability and similar "
                     'undocumented parameters',
                     'Implement stricter input validation for AI prompts',
                     'Enhance monitoring of AI interactions for suspicious '
                     'activity',
                     'Conduct security audits of AI systems with access to '
                     'sensitive data',
                     'Educate users on risks of AI prompt manipulation'],
 'references': [{'source': 'Varonis Threat Labs'}],
 'response': {'containment_measures': 'Microsoft planned a patch and CVE '
                                      'assignment',
              'remediation_measures': 'Disabling vulnerable URL parameters, '
                                      "patching the 'autorun=1' flaw"},
 'threat_actor': 'Varonis Threat Labs (Researchers)',
 'title': "Microsoft Copilot Vulnerability 'CoSnitch' Exposed via AI "
          'Self-Disclosure',
 'type': 'AI Vulnerability Exploitation',
 'vulnerability_exploited': "Undocumented 'autorun=1' parameter, lack of "
                            'separation between data and system instructions '
                            'in LLMs'}
Great! Next, complete checkout for full access to Rankiteo Blog.
Welcome back! You've successfully signed in.
You've successfully subscribed to Rankiteo Blog.
Success! Your account is fully activated, you now have access to all content.
Success! Your billing info has been updated.
Your billing was not updated.