A significant data security vulnerability has been identified in xAI’s Grok Build coding command-line interface (CLI), version 0.2.93, which was found to be indiscriminately uploading entire Git repositories, complete with full commit histories and potentially sensitive unredacted credentials, to a Google Cloud Storage bucket managed by xAI. This discovery, made by an independent security researcher publishing under the pseudonym cereblab, revealed a profound discrepancy between the data required for the AI model’s operation and the vast quantities of data being exfiltrated, raising serious concerns about user privacy, proprietary code protection, and the broader implications for trust in AI-powered development tools. The incident underscores a critical lapse in data handling practices, where default settings allowed for broad data collection far exceeding functional necessity, even when user-facing privacy controls were ostensibly engaged.
The Unveiling: A Deep Dive into Data Exfiltration
The vulnerability came to light through meticulous testing by cereblab, who specifically examined Grok Build CLI version 0.2.93. The researcher’s methodology involved intercepting network requests generated by the tool during a coding task. What they uncovered was alarming: instead of uploading only the specific files pertinent to a given coding task, Grok Build was transmitting the entirety of local Git repositories, encompassing every file tracked by Git and, crucially, the full, chronological commit history of the project. This extensive data dump was directed to a Google Cloud Storage bucket operated by xAI, identified as grok-code-session-traces.
To validate the findings, cereblab performed a crucial test. A specific file, src/_probe/never_read_canary.txt, containing a unique marker, was deliberately planted within a test repository. This file was then explicitly instructed not to be opened or processed by the Grok agent. Despite this clear instruction, analysis of the intercepted upload revealed that the entire Git bundle, including this "never-read" canary file, was transmitted. The researcher was able to clone the git bundle directly from the intercepted request, successfully recovering the canary file verbatim along with the repository’s complete commit history. This same test was replicated on a second, entirely unrelated repository, yielding identical results, thus establishing a consistent pattern of comprehensive data exfiltration.
The scale of this data collection was particularly stark. The upload mechanism operated on a separate network channel from the actual model interaction. While the model’s direct communication to /v1/responses for processing a coding task amounted to a mere 192 KB on a 12 GB repository, the storage channel to /v1/storage moved an astonishing 5.10 GiB of data. This represented a staggering 27,800-fold difference between the data the model genuinely needed for its task and the volume of data that departed the user’s machine. The large repository upload was broken down into 73 chunks, each approximately 75 MB in size, with every chunk returning an HTTP 200 success code, confirming the successful and complete transmission of the repository data. The destination bucket, grok-code-session-traces, was explicitly referenced within the CLI’s binary and in a staged metadata.json file, further corroborating the intended storage location. It is important to clarify that cereblab‘s findings definitively established the transmission, acceptance, and storage of this data, not necessarily its immediate training or direct access by xAI staff. However, the mere fact of its unauthorized collection and storage constituted a severe security breach.
The Peril of Unredacted Secrets and Misleading Controls
Beyond the wholesale upload of repository data, cereblab uncovered an even more critical vulnerability concerning sensitive credentials. When Grok Build did read a file as part of its operational task, its contents were fed into the model’s processing turn. Crucially, if this file was a tracked .env file containing environment variables—often used for storing API keys, database passwords, and other sensitive configurations—its contents, including canary API_KEY and DB_PASSWORD values, were transmitted unredacted. These unredacted secrets were not only sent to the model but also subsequently landed in a session_state archive, which was then bound for the same cloud storage bucket. While the secrets used in cereblab‘s tests were fake, preventing any real-world data leak in this specific instance, the inherent behavior of transmitting and storing unredacted credential files posed an immense risk.
Adding to the gravity of the situation was the inadequacy of user-facing privacy controls. Most developers, when using such tools, would instinctively look for options to limit data sharing. Grok Build offered a setting labeled "Improve the model," which users might reasonably assume would govern the extent of data collection. However, even with this setting explicitly turned off, Grok Build continued its comprehensive repository uploads. Further investigation revealed that the server’s own /v1/settings response continued to return trace_upload_enabled: true, effectively overriding any user-side attempt to disable data collection. This highlighted a fundamental design flaw: the "Improve the model" toggle primarily governed whether user data would be used for model training. It did not, however, control whether the user’s source code and sensitive files would leave their machine entirely. These were two distinct control mechanisms, and only one, the less critical one in terms of immediate data exfiltration, was exposed and controllable by the user.
Industry Context and Grok Build as an Outlier
The landscape of AI-powered coding agents inherently involves some level of source code transmission to remote models for processing tasks. This fundamental requirement means that a "local only" mental model is generally inappropriate for these tools. However, the scope of data collection exhibited by Grok Build far exceeded industry norms and reasonable expectations.

A comparative analysis conducted by cereblab across various leading AI coding tools underscored Grok Build’s anomalous behavior. Competitors like Claude Code and Codex were observed to send no repository bundle whatsoever during their operations. Google’s Gemini, in an idle test, also sent no such bundle, though its realistic-task run was quota-blocked before completion, preventing a full comparison. Grok Build, in stark contrast, was the definitive outlier due to its "wholesale collection of the workspace."
The implications of such comprehensive data exfiltration are profound. A typical Git repository can contain a wealth of sensitive information:
- Proprietary Code: The core intellectual property of companies, which, if exposed, could lead to competitive disadvantages or direct theft.
- Internal URLs and Network Configurations: Information about internal network structures, staging environments, and unlisted endpoints that could aid malicious actors in reconnaissance or targeted attacks.
- Customer Data: In some cases, sample data or test data that might inadvertently contain anonymized or even real customer information.
- Credentials in Commit History: Perhaps the most insidious risk. Even if a developer has diligently removed sensitive API keys, database passwords, or other credentials from the current working tree, these often persist within the repository’s commit history. A full repository upload, including history, would expose these previously committed secrets, making them vulnerable even if they are no longer actively present in the project’s latest version.
The broad boundary of sending an entire tracked repository and its history, rather than just the task-relevant files, represented a severe deviation from best practices for data minimization and security in cloud-based development environments.
xAI’s Swift, Server-Side Remediation and Communication Gaps
Following the public disclosure of cereblab‘s findings, xAI initiated a swift, albeit unannounced, remediation. On July 13, just days after the vulnerability was identified, cereblab retested the same 0.2.93 binary of Grok Build CLI and observed a complete cessation of /v1/storage upload requests. Six subsequent retests consistently showed zero storage uploads. Crucially, the server’s response to configuration queries had also changed, now returning disable_codebase_upload: true and trace_upload_enabled: false. This indicated that the fix was implemented via a server-side switch, altering the behavior of existing client binaries without requiring a new software update. This was further corroborated by developer Peter Dedene, who reported seeing the same flag changes for his own account, suggesting the fix was rolled out more broadly than just to cereblab‘s test environment.
However, xAI’s communication regarding this critical security issue has been notably informal and lacking in standard protocols. Instead of issuing a formal security advisory, a comprehensive changelog note, or a dedicated press release, xAI addressed the matter primarily through posts on X (formerly Twitter). The official @SpaceXAI account stated that enterprise teams operating under a Zero Data Retention (ZDR) policy never had code or trace data stored, and that API-key usage also respected ZDR. For individual consumers who had not enabled ZDR, the account advised running /privacy in the CLI to disable retention and delete previously synced data. Elon Musk, CEO of xAI, subsequently amplified this message, stating unequivocally that all user data uploaded prior to the fix would be "completely and utterly deleted," with "nothing left behind."
While these statements from xAI leadership are reassuring, the method of communication—via social media rather than official security channels—leaves several questions unanswered and creates ambiguity. It is unclear whether the server-side fix reaches every account or is a permanent solution. The lack of a formal public record or detailed technical explanation raises concerns about transparency and accountability, particularly for an AI company handling sensitive developer data.
Immediate Recommendations and Lingering Concerns
For any developer who has previously used Grok Build, the immediate and most critical action is to rotate any credentials that the tool could have potentially accessed or transmitted. This includes:
- Anything Grok read: Any credential file directly opened or processed by the agent during a task.
- Anything in a tracked file: Any sensitive data present in files that were part of the Git repository and thus part of the comprehensive upload.
- Anything in the Git history: Critically, any secrets that were committed to the repository at any point, even if they were subsequently deleted from the working tree or from later commits. The full commit history carried by the bundle would have preserved these past secrets.
It is important to understand the distinction: a file that was gitignored and never committed to the repository would have stayed out of the uploaded bundle. However, a file that was committed, even if deleted later, would have ridden along in the history and is therefore compromised.

The server-side nature of xAI’s fix also introduces a layer of vulnerability. A separate analysis of build 0.2.99 (a slightly newer version) found that the underlying upload code for repositories still existed within the binary, merely held in abeyance by the server-side flag disable_codebase_upload: true. This means that xAI theoretically retains the technical capability to re-enable the full repository upload feature without requiring users to update their client software, a possibility that could undermine trust and create future security risks if not managed with extreme transparency.
Crucially, xAI has yet to provide comprehensive answers to several fundamental questions:
- Why was the default behavior to upload entire Git repositories, including full commit history, in the first place?
- How long was this extensive user data retained on xAI’s cloud storage?
- How many users were affected by this default data collection policy?
The incident serves as a stark reminder that a user-facing "training opt-out" is not a guarantee that sensitive code or data will remain on the local machine. Developers must exercise vigilance and actively verify what data their AI-powered tools are transmitting.
Broader Implications for AI Development, Trust, and Regulation
This incident with xAI’s Grok Build has far-reaching implications for the burgeoning field of AI-assisted development and the broader tech industry. It highlights a critical tension between the desire to rapidly develop and improve AI models—which often benefits from vast amounts of data—and the fundamental principles of user privacy, data security, and intellectual property protection.
The lack of granular, transparent controls for data transmission represents a significant trust deficit. Developers, who are increasingly relying on AI tools to augment their workflow, need assurances that their most valuable assets—their code and credentials—are handled with the utmost care and only collected with explicit, informed consent for clearly defined purposes. When tools silently exfiltrate entire repositories, regardless of the stated purpose, it erodes this trust.
This event also underscores the vital role of independent security researchers like cereblab. Their proactive testing and transparent disclosure are essential for uncovering hidden data practices and holding AI companies accountable. Without such scrutiny, widespread vulnerabilities could persist unnoticed, leading to significant harm.
From a regulatory perspective, incidents like this could fuel calls for stricter data governance frameworks specifically tailored for AI systems. Existing regulations like GDPR and CCPA already mandate transparency and user control over personal data, but the unique nature of code and proprietary information in AI contexts may necessitate further clarification and enforcement. Companies developing AI tools must adopt a "security by design" and "privacy by design" approach from the outset, embedding robust data minimization strategies, transparent data flows, and explicit user consent mechanisms into their products.
Ultimately, the xAI Grok Build incident is a potent case study in the challenges of balancing innovation with responsibility in the age of AI. It sends a clear message to both AI developers and users: vigilance, transparency, and explicit control over data are not optional luxuries but fundamental requirements for building a secure and trustworthy AI ecosystem. The onus is on companies to provide clear, actionable privacy controls and transparently communicate their data handling practices, while users must remain critically aware of the data footprint of the tools they integrate into their workflows.
