Aaron Swartz faced 35 years for 70GB. Meta scraped 80TB for AI with barely a scratch.

Aaron Swartz faced 35 years for 70GB. Meta scraped 80TB for AI with barely a scratch.

The legal system threatened Aaron Swartz with 35 years in prison for downloading 70 gigabytes of academic papers, while Meta torrented 80 terabytes of books for AI with almost no consequences.

A post trending on Hacker News highlights tech's stark double standard around data scraping. RSS co-creator Aaron Swartz was aggressively prosecuted for downloading academic papers from JSTOR to make knowledge public. Prosecutors hit him with asset forfeiture, $1 million in fines, and decades in prison before he took his own life. Meanwhile, Meta torrented 80 terabytes of books to train its proprietary AI models—a corporate operation likely to end in a minor financial slap on the wrist while its models print money.

Why it matters: The contrast shows who web scraping laws actually protect. Scrape data as an individual to make knowledge accessible, and the government treats you as a criminal. Scrape data at 1,000 times that scale to power proprietary commercial models, and it is just a line item on a corporate balance sheet.

Know this: Swartz collected 70 gigabytes of papers to archive and share research. Meta harvested 80 terabytes of copyrighted books to train proprietary systems that enrich its corporate owners.

Turns out the legal risk of downloading the open web depends entirely on whether you can afford to write off court cases as a business expense.