Since early 2023, photographer Jingna Zhang and a small crew of volunteers have kept an image-sharing and portfolio app called Cara running. It has attracted about 1.5 million artists. What drew them was a shared opposition to the unauthorised use of their work to train AI models, and a desire to show their art without being exploited by Big Tech.

But while Cara filters out AI images and offers some minimal protective features, preventing scrapes themselves is nearly impossible. Beginning on August 13, the platform was subjected to three major scrapes. Server fees spiked, and creators who had migrated from platforms like Instagram — where content is explicitly available to Meta as training data — were alarmed.

Three scrapes

  • 12 million works. The person behind the first incident announced on a Reddit community that he had obtained a 12-terabyte archive — more or less Cara's entire library of publicly available images. In his since-deleted post he called it "a fun project" and said it cost him less than $10.
  • 8.5 million links. A second scraper uploaded the links, along with metadata such as usernames and tags, to Hugging Face. Responding to takedown requests, the platform said it would notify the user to remove the personal metadata but could not touch the URLs, since "no copies of the artworks are hosted here" and the links point to "the copies the artists published on Cara."
  • 123,000 images. On August 22 a third scraper put those, along with text posts and user bios containing personal information, on a site called Academic Torrents.

Zhang found out when users tagged her: the scraper was "gloating and looking for other people to join him to do something with the dataset on Reddit." Her assessment is short: "I just feel it's targeted and very hurtful." She adds that laws have not caught up on protections against such data harvests, which means scrapers can often justify the act as technically legal.

The scraper changed his mind

Here is the unexpected part: the person who grabbed all the art came to regret the stunt and agreed to collaborate with Zhang on a new open-source tool to protect artists.

Referred to only by the screen name "Heft" because of doxing and death threats he says he received, he is a student in North America with a software background and an interest in digital preservation and archival projects. He tells Wired that scraping Cara was originally nothing more than a technical project and that he had no intention of making the data public.

"In retrospect, not only deliberately targeting Cara but presenting it the way I did in the post was cruel and thoughtless," he says. "I missed the consequences that this would have beyond causing a bit of anger." He saw people "sharing how they were having panic attacks over the scrape, how they deleted their entire portfolios from the internet."

Zhang's bind

Zhang is frustrated not only by the scraping but by the confusion around what her team can realistically do. Some users have already deleted their portfolios and left. "We have done the right things within limits without making it horrible to use," she says of the current safeguards, including temporary measures like login gates — which, she adds, are not really a solution to an ongoing, internet-wide problem.

What she says to those leaving is striking: "If it makes them feel better, deleting your work and leaving Cara, I support that. But I don't want to give people the misconception that if they go somewhere else, they are safer, because they're not. Bigger platforms get scraped more, so that makes me feel worse."

Zhang launched a fundraiser for legal fees with a goal of $120,000 and has raised more than $100,000. She is separately part of two class actions brought by visual artists: one against Stability AI and Midjourney, the other against Google.