He Scraped a Million Artists' Work for Fun. Now He's Helping Them Fight Back.
A student who downloaded Cara's entire image library and bragged about it online ended up regretting the stunt and building a protection tool with the platform's founder.

Key points
- Starting August 13, 2025, Cara, an artist portfolio platform, suffered three separate mass downloads of its images by outside parties.
- The first scraper downloaded roughly 12 million images, nearly Cara's entire public library, for less than $10 in server costs.
- Cara founder Jingna Zhang has raised more than $100,000 toward a $120,000 legal fund to explore copyright and cybersecurity options.
- The first scraper later apologised, deleted his copy of the data, and is now co-building a free notification tool called Lantern with Zhang.
- Lantern scans public AI training datasets and alerts artists if their work appears in one, so they can request removal.
Jingna Zhang did not set out to run a tech company. She is a photographer who, along with a small group of volunteers, built Cara, a social media and portfolio app designed for artists who want to share their work without handing it to big platforms as free fuel for AI training. By this summer, about 1.5 million artists had joined.
Then, in the space of ten days, three separate people downloaded huge chunks of it.
What exactly happened to Cara?
Three bulk downloads of Cara images, known as scrapes, hit the platform in quick succession starting August 13. The first and largest came from a person posting on Reddit under the name MandarinDawnPoppy994, who published a 12-terabyte archive, a file size roughly equal to three million novels, containing 12 million images and bragged that the whole operation cost him less than $10. Zhang found out because her users spotted the post and tagged Cara directly.
A second person pulled 8.5 million image links plus user metadata, including usernames and tags, and uploaded the collection to Hugging Face, a platform where AI developers share data and software tools. Hugging Face told Zhang's team it would ask the user to remove the personal metadata but would not touch the image links, since the links pointed back to Cara's own servers and no copies sat on Hugging Face's systems. A third scraper took 123,000 images along with user bios containing personal details and posted everything to a file-sharing site called Academic Torrents.
Cara already offered some defences. It integrates Glaze, a tool that subtly alters image files so that AI training software misreads the visual style, making it harder for a model to copy an artist's look. But Glaze cannot stop someone from simply downloading the files. As Zhang put it in comments first reported by Wired AI: "Laws are not caught up on" protections against data harvests of this kind, and scrapers can often argue the act is technically legal.
Who was the first scraper, and what changed?
The person behind the Reddit post goes by the screen name "Heft" and is a software student in North America with an interest in digital archiving. He told Wired AI he originally scraped Cara as a personal technical project, with no plan to publish the data, but made what he called "a foolish decision to attempt to ragebait with the dataset on Reddit."
The reaction shook him. Artists described panic attacks. People deleted entire online portfolios built over years. Heft reached out, apologised, and deleted his copy of the archive.
He then joined Cara's Discord server as a volunteer troubleshooter, explaining to Zhang why proposed fixes, such as login gates, would not hold for long. "He's just helping us, taking the time to explain" the structural weak spots, Zhang said.
The two are now building Lantern together. The tool creates a one-way fingerprint for each image, meaning it can recognise a picture without storing a copy of it. Lantern then scans publicly available AI training datasets on a regular basis. If it finds a match, it alerts the artist and links them to the dataset so they can file a takedown request.
Lantern is not a complete solution. Heft is clear that no website can be made truly impossible to scrape. But it gives artists something they currently lack: a way to know when their work has already left the building.
What should artists on Cara do right now?
Zhang says artists who leave for another platform are not necessarily safer. Larger platforms get scraped more often, not less. If leaving makes someone feel better, she supports that choice, but she does not want people to assume a different site offers real protection.
Practically speaking: watch for updates on Lantern, which is still being built. If your work matters to you, consider using Glaze before uploading anywhere public. And if you spot your images in an AI dataset, a formal takedown notice under copyright law is currently one of the few concrete tools available.
Zhang's GoFundMe for legal fees had raised more than $100,000 toward its $120,000 goal as of late August 2025.
Common questions
Is scraping images from a website illegal?
Not automatically. Downloading publicly visible images may be technically legal in many places even without the creator's permission, which is why Zhang says the law has not kept pace with how AI companies and individuals collect training data.
What is a training dataset, and why does it matter to artists?
A training dataset is a large collection of images, text, or other content used to teach an AI model what things look like or how language works. If an artist's work ends up in one, the AI can learn to imitate their style, often without the artist knowing or being paid.
Will Lantern actually protect artists?
Lantern cannot stop someone from downloading images in the first place. What it can do is tell artists where their work has ended up, so they have a fighting chance to request removal before the data spreads further.



