ai.type Data Breach (2017): What Was Exposed & What To Do
SourceBreach data provided in part by Have I Been Pwned, used under CC BY 4.0.
The ai.type Data Breach (2017) (reported December 5, 2017) exposed Address book contacts, Apps installed on devices, Cellular network names and Dates of birth belonging to roughly 20.6M people. If you have an account with them, your information may now be circulating on the open web and with data brokers. Here’s exactly what happened, how to check if you were affected, and what to do next.
What happened
The breach was reported on December 5, 2017. Investigators found that ai.type had stored a substantial volume of user-related records in a MongoDB deployment that lacked access controls. The exposed collection measured 577 gigabytes and included records tied to roughly 20.6 million people. Discovery occurred through external research rather than through disclosure by the company itself. No information has been released about the duration the instance remained open or the precise method by which the data first became reachable.
How a breach like this happens
Incidents involving publicly accessible databases often begin with default or misconfigured network settings on database servers. MongoDB instances, when installed without authentication requirements or without restrictions on inbound connections, can be reached by any party that locates the correct IP address and port. Once identified through automated scanning, such servers allow direct queries or bulk exports of stored collections. The absence of encryption or logging can further delay detection by the data owner.
ai.type and its sector
ai.type developed a virtual keyboard application used on mobile devices. Applications of this type routinely collect and store user inputs, device identifiers, and contact information to provide features such as predictive text and personalization. The sector handles data generated during everyday device use, including details that can link individuals to their social and professional networks. Exposure of such records therefore carries implications for both individual privacy and downstream uses of the information by third parties.
What data was at risk
The dataset contained several categories of information. The following items were identified in the exposed records:
- Address book contacts
- Apps installed on devices
- Cellular network names
- Dates of birth
- Device information
- Email addresses
- Genders
- Geographic locations
Additional references to social media profiles appeared in reporting on the incident. The exact scope of every field present in the 577-gigabyte collection remains unconfirmed beyond the categories listed above.
Why it matters
Records that combine email addresses with device details, locations, and contact lists can support targeted follow-on activity such as phishing or account enumeration. For the organization, the incident highlighted the operational consequences of storing production data without network-level protections. Individuals whose information appeared in the dataset received no direct notification from ai.type, leaving them to learn of potential exposure through secondary sources or breach-notification services.
Were you affected?
Individuals can check whether their email address appears in known breach datasets by using publicly available lookup services that index records from incidents such as this one. Practical next steps include reviewing account passwords for any services tied to the exposed email addresses and enabling multi-factor authentication where available. Monitoring for unusual login attempts or unsolicited messages that reference personal details from the records can also help limit further impact.
AICompiled with AI assistance from public sources and published under our editorial standards.
How this breach connects
More recent breaches
Netshoes Data Breach (2017)Open CS:GO Data Breach (2017)B2B USA Businesses Data Breach (2017)Exposed VINs Data Breach (2017)Latest breaches
Read GalaxyWarden’s full analysis of the ai.type Data Breach (2017) →
Verified breach. Breach data provided in part by Have I Been Pwned, used under CC BY 4.0.
Breach listings — particularly those originating from ransomware or leak sites — are third-party claims that may be unverified, incomplete, or inaccurate. A listing does not by itself confirm that a breach occurred or that any specific data was exposed. Severity is an automated assessment, not a definitive rating. Verification status is shown where available.
Attributions to threat groups and methods reflect public reporting and, in some cases, unverified claims made by the groups themselves; they may be incomplete or later revised. Recent Breaches and GalaxyWarden are independent and are not affiliated with, and do not endorse, any company or group named on this page. This information is aggregated from public sources for awareness only and is not legal, security, or investment advice.