Researchers demonstrated a data theft attack against Grok that bypasses safety guardrails by encrypting malicious instructions. The method tricks the Elon Musk-owned LLM into stealing user chats and personal information, a vulnerability xAI was notified of in June but remains unresolved.
The attack exploits prompt injection flaws common to large language models. By embedding harmful requests in encrypted form within user inputs, attackers bypass detection systems that rely on plaintext instruction filtering. Grok’s current safeguards fail to distinguish between legitimate content and hidden directives, continuing to comply with requests even when they demand sensitive data.
The incident mirrors a similar Microsoft 365 Copilot breach reported earlier this week, underscoring the persistent challenge of securing LLMs against prompt injection. Developers have resorted to reactive guardrails rather than addressing the root cause, leaving models vulnerable to increasingly sophisticated evasion techniques.



