
How to Profit from Selling Employee Data After a Company Goes Bankrupt

By Sleepy @sleepy0x13
The AI era has spawned a new industry chain: acquiring internal data left behind by defunct companies.
Emails, chat logs, project documentation, and work tickets used to be just digital debris awaiting deletion after a company shut down. Now, they are being revalued, packaged, sold, and fed directly into the training pipelines of AI companies.
Morticians preserve a person's final dignity after death. For a dying corporation, dignity lies in proving that what it leaves behind still holds value.
On August 17, at a bankruptcy asset auction, Google bid $10 million to acquire all enterprise data belonging to Spirit Airlines. There was one other bidder, Mercor, which offered $7.5 million—falling short by $2.5 million.
The lot was divided into three parts. First, approximately 100 million employee emails. Second, 500 million Microsoft Teams messages. Together, the two totaled 600 million records. Third, calendars, spreadsheets, financial databases, project files, operational logs, and a suite of internal software. Passenger profiles, frequent flyer data, and similar records were not included.
Six hundred million messages: if a person spoke 100 sentences a day, it would take them over 16,000 years of non-stop talking.
That works out to roughly 1.67 cents per message. Americans call the one-cent coin a penny—something people hardly bother to pick up off the ground. In other words, a single sentence spoken by a Spirit employee on Teams was priced at about a penny and a half.
Rewind to 1980. Spirit Airlines was born in Detroit, originating as a trucking business called Charter One before pivoting to aviation. In 1992, it renamed itself Spirit. Over the next 34 years, it pioneered the ultra-low-cost carrier (ULCC) model across the United States, becoming a benchmark copied across the industry.
It once had a solid foundation: an all-Airbus fleet of 205 aircraft, around 300 flights a day, and revenue of roughly $5 billion in 2024. Yet that same year, it suffered a net loss of around $1.2 billion, ultimately entering bankruptcy burdened by about $9 billion in debt.
Late one night in May 2026, the company announced it was halting operations. The next day, approximately 17,000 employees learned from the news that they were out of a job.
The transaction is still going through legal procedures and awaits approval from a bankruptcy judge. Spirit filed for Chapter 11 liquidation; with no bankruptcy trustee appointed, the company remains under court supervision, selling off its remaining assets piecemeal. Sensitive data such as employee emails and Teams messages must first be handled by an independent third party to scrub personally identifiable information (PII) like names and email addresses. This firm will be selected by Google, with all costs borne by Google.
Furthermore, across public records, there is virtually no precedent for a bankrupt company selling its internal communications data to an AI company. This deal is very likely the first of its kind in history.
Veteran US tech outlet Gizmodo headlined the story: Spirit is dead, but its ghost will haunt Google's servers for generations.
A New Business
There have always been people who clean up after dead companies. Lawyers, liquidators, and auction houses have been doing it for decades. Aircraft, office furniture, trademarks, and patents—anything that could be sold was sold long ago.
What is genuinely new this year is that internal employee data has also been placed on the auction block.
This business has emerged driven by two main factors.
First, more companies are dying.
In the first quarter of 2024, the failure rate for US startups surged by 58% year-over-year, while the number of active venture capital firms dropped by 62% from its peak.
Capital didn't necessarily dry up; it simply became heavily concentrated in AI. In 2024, US AI startups raised a record-breaking $97 billion. There is still plenty of money in the capital markets, but investors are increasingly reluctant to deploy it anywhere outside of AI.
As a result, a wave of companies that previously could have survived on venture rounds hit a wall much sooner. In August 2024, a16z-backed fintech firm Tally shut down. Despite raising a total of $172 million, reaching a peak valuation of $855 million, and progressing to Series D, it was unable to secure its next round of funding.
Second, data has become expensive.
The data consumption required for training large language models has reached staggering proportions. GPT-4 was trained on roughly 13 trillion tokens. For comparison, Google Books scanned around 40 million books over four decades, translating to about 4 trillion tokens. That means a single training run of GPT-4 consumed more than three times the volume of Google Books.

Epoch AI calculated that high-quality human language data—books, news, Wikipedia—could be virtually exhausted around 2026. High-quality text available across the public web for training is dwindling rapidly, leading to a rising share of synthetic data. Going forward, AI companies searching for fresh data have little choice but to venture beyond the public web. Years of accumulated internal emails, chat logs, and workplace documents have entered their crosshairs.
Meanwhile, researchers project the market for AI training datasets will reach $9.7 billion by 2030. When factoring in all forms of licensing, the total addressable market could reach $67.5 billion.
On one side, a growing number of companies are shutting down and liquidating, leaving behind vast troves of previously unpriced internal data. On the other side, AI companies have a voracious appetite for non-public web data. The convergence of these two trends has created, for the first time, the conditions for internal enterprise data to be traded at scale.
What makes this data truly valuable is the rise of enterprise AI agents. Gartner predicts that by 2026, 40% of enterprise applications will embed task-oriented AI agents, up from less than 5% just a year earlier.
The training materials required for enterprise agents differ significantly from those for general-purpose LLMs. Public web pages provide knowledge, natural language, and finished output, but they rarely capture the messy, authentic workflow of a company. Real-world communication and collaboration involve intricate details: how a requirement is framed, how team members debate it, how tasks are delegated, how bugs are rectified, and how final deliverables are shipped.
These entire processes are preserved inside corporate emails, chat logs, support tickets, and project documents.
Once clear buyers and use cases emerged for this data, material that was previously deleted upon shutdown gained distinct standalone value on the market.
The Corpse Handlers
Liquidation lawyers used to handle this work alone; now, three new groups have joined the fray.
The first group consists of dissolution service providers, exemplified by SimpleClosure.

This company does only one thing: help startups die with dignity. Founded in 2023 with a $1.5 million pre-seed round, it raised a $15 million Series A in May 2025 led by TTV Capital. Even Carta, which manages equity and corporate administration for countless US startups, shut down its own wind-down services to invest in SimpleClosure and route its client requests there.
By October 2025, SimpleClosure had handled "funerals" for over 1,000 companies. Crunchbase dubbed it "A Better Way To Fail." American founders often preach "fail fast." SimpleClosure argues that failing fast isn't enough—it must also be dignified. Its website even features a pricing calculator: input your company's parameters, and it estimates the cost of shutting down.
These wind-downs are not done for free. On April 16, 2026, SimpleClosure launched Asset Hub, dedicated to monetizing intangible assets left behind by shuttered startups. Alongside brands, software, and customer lists, internal operational data like Slack logs, emails, and Jira tickets was officially placed on the shelf for the first time.
This shows that even before Spirit Airlines, the market had begun experimenting with pricing defunct companies' internal data—albeit among smaller startups via private brokering.
There is already a concrete case. When transcription and captioning company cielo24 shut down, it sold 13 years of accumulated Slack messages, internal emails, and Jira tickets through SimpleClosure. CEO Shanna Johnson later told Forbes that the dataset fetched several hundred thousand dollars.
For a company that had decided to close its doors, data that once needed cleanup became an asset recovered during liquidation.
SimpleClosure offloaded the data-selling component to Protege. Protege is a data marketplace specializing in AI training data licensing, having raised $30 million in January 2026 led by a16z, founded by Bobby Samuels. Protege initially entered the market through medical imaging, sourcing millions of scans for pre-training buyers within 30 days.
Now, Protege is applying its data licensing and transaction infrastructure to the internal communications of defunct companies. SimpleClosure handles company wind-downs and asset collation, while Protege identifies buyers and executes data licensing deals.
The second group of participants consists of bankruptcy courts and liquidation attorneys. For decades, they cataloged planes, desks, chairs, trademarks, and patents. Today, mailboxes, Slack logs, and other internal data are appearing on their asset inventories.
Under US bankruptcy law, this data can be treated as intangible property of the bankruptcy estate and sold under court supervision. Law firms have started setting up specialized practices for data preservation, forensics, and organization in insolvency cases. Redgrave LLP, for instance, runs dedicated restructuring and forensics practices.
The third group comprises technical e-discovery service providers. Firms like KLDiscovery, Epiq, and Consilio routinely collect, organize, host, and review enterprise data. Content from email inboxes, Teams, and SharePoint must first be exported, archived, and packaged by them into formats ready for transaction pipelines.
The e-discovery sector already has well-established fee structures. EDRM regularly publishes pricing surveys, covering standard billing units such as collection per gigabyte, monthly hosting per gigabyte, and document review costs.
Turning Spirit's 600 million messages into an auction lot required extensive foundational work by these providers. Yet unlike court filings or auction bids, this technical layer rarely appears in news coverage; outsiders seldom see who handles the data or how it is processed.
The Autopsy Checklist
For enterprise AI agents, training requires more than just encyclopedic facts or standardized answers; it requires learning the judgment, collaboration, debugging, and execution workflows of real workplace environments.
Finished products show the model what was built; internal communications show how it was actually built.
This shift is already evident in agent training datasets. Previously, a code training sample might have been a few hundred tokens showing a diff of a few lines of code. Today, an agent training sample often encompasses the end-to-end process: requirement comprehension, file localization, code modification, and test verification.
Training data is evolving from standalone answers to complete task execution logs. For an AI agent, the end result matters, but the true training value lies in the human judgment, operational steps, and iterative feedback generated along the way.
Compared to operating businesses, defunct companies offer data that is far easier to trade. While a company is operational, selling internal communications triggers trade secret concerns, employee privacy disputes, non-compete issues, and customer relation risks, making legal and executive teams exceptionally cautious. In liquidation, however, the sole objective shifts to maximizing residual asset recovery for creditors.
That is why the exact same internal data—virtually unsellable while a company is alive—can be reappraised as a valuable asset upon its death.
The enterprise software industry has long recognized the value of this data. Salesforce has consistently viewed corporate communications inside Slack as a vital data asset, while Microsoft CEO Satya Nadella has repeatedly emphasized that when enterprises deploy AI, the true differentiator is their proprietary data and working context.
Historically, this data served only the enterprise itself. Now, as AI developers actively hunt for internal corporate records, it has found clear outside buyers for the very first time.
Once buyers exist, the next question is pricing.
The Art of Pricing
Although this marketplace is still in its infancy, several pricing benchmarks have started to emerge.
For defunct startups handled by SimpleClosure and Protege, individual transactions generally range between $10,000 and $100,000. Mercor has offered up to $300,000 for employee chat logs and emails from acquired startups.

Spirit pushed the price tag straight into eight-figure territory. Mercor bid $7.5 million, and Google ultimately offered $10 million, competing for approximately 600 million internal communications and related enterprise records. Calculated solely on those 600 million messages, the average price comes to about 1.67 cents per message.
This unit price is not particularly high. Reuters reported in 2024 that Photobucket negotiated licensing terms with AI companies for roughly 13 billion photos and videos, asking 5 cents to $1 per photo and over $1 per video. In B2B data marketplaces, a single contact record can sell from several cents to several dollars depending on completeness and accuracy.
Yet these figures have not coalesced into a standardized benchmark. The true value of a defunct company's internal data is currently negotiated on a case-by-case basis. Total volume, industry sector, time span, completeness, uniqueness, and the buyer's intended use case all influence final valuation. Spirit's $10 million deal serves primarily as one of the few visible large-scale reference points.
When compared to mature data licensing markets, the difference is stark. Reddit licenses user posts and comments to Google for roughly $60 million annually; News Corp inked a five-year content licensing deal with OpenAI worth approximately $250 million; xAI's partnership with Telegram reached $300 million; and Apple's image licensing deals with Shutterstock range between $25 million and $50 million.
Those markets benefit from established buyers, licensing models, and pricing precedents. The trading of defunct enterprise internal data is just getting started, with no codified rules yet on which data is most valuable, or whether it should be valued per record, per gigabyte, or per package.
Framed against Spirit's own operational scale, $10 million takes on another perspective. In 2024, Spirit generated nearly $5 billion in annual revenue—averaging roughly $13.7 million per day. Google acquired this entire data corpus for less than a single day of the airline's normal operating revenue.
For the bankruptcy estate, it was merely one piece of residual asset recovery; for the AI giant, it secured a treasure trove of non-public data capturing the authentic workings of a major enterprise.
Scrubbing and Handover
Once a deal is struck, the data cannot simply be handed over to the buyer. The transfer process typically involves several stages: export, de-identification, packaging, court approval, and final closing.
The first step is export. Slack enterprise data is typically exported as ZIP archives containing JSON files organized by channel, along with user profiles and attachments.
Microsoft 365 environments utilize forensic tools to export emails and Teams messages. Spirit's deal involves roughly 600 million communications—a massive dataset requiring batched processing. Public filings have not disclosed whether this is being handled by an e-discovery vendor or the buyer's engineering team.
The next step is de-identification: stripping out information that points directly to specific employees. Common methods include identifying and redacting personal details like names, phone numbers, and addresses, or replacing sensitive fields with anonymized identifiers. In certain statistical and training contexts, random noise perturbation is added to further minimize re-identification risk.
However, de-identification does not guarantee absolute anonymity. Even when names and email addresses are removed, if the text retains sufficient context regarding roles, timestamps, locations, or behavioral patterns, individuals can still potentially be re-identified through cross-referencing.
That makes the entity responsible for this step critical in the Spirit transaction. Under current terms, data de-identification will be managed by an independent third party selected and paid for by Google. Google has also pledged not to use the dataset to re-identify any individuals.
Selling data out of bankruptcy is not entirely new. Over the past two decades, companies like Toysmart, Borders, RadioShack, and 23andMe all navigated the disposal or sale of customer data during insolvency. Because those cases involved consumer privacy, they faced rigorous scrutiny from courts, regulators, and state attorneys general.
Spirit differs in that the core of the sale is not passenger manifests, but employee emails, Teams messages, and internal workplace files. Existing bankruptcy frameworks lack well-defined rules for handling this category of employee data compared to consumer records.
Controversy has already erupted. Spirit's flight attendants union filed formal objections to the sale, prompting the court to delay approval. What began as a routine asset sale has intersected with labor rights and the legal boundaries of AI training.
If final court approval is granted, the data will move to closing. Yet public filings reveal very little about the path from packaged data to Google's internal infrastructure. Exactly how the data will be transferred, whether additional cleaning will occur, and in what format it will enter product development or model training remains undisclosed.
Only when that process concludes does the data truly complete its transition from bankruptcy scrap into an AI asset.
The Buyers
Currently, the most prominent buyers for this data are frontier model developers like Google and training data providers like Mercor.
Google's official rationale for the Spirit transaction cites product improvement and AI development. For Google, the value of this corpus lies in its documentation of real enterprise operations: shift scheduling, cross-team coordination, internal chatter, project management, issue escalation, and executive decision-making.
These nuances cannot be scraped from the public internet. For enterprise AI agents in particular, what matters most is not merely the final output, but how complex tasks are navigated and resolved within a corporate hierarchy. Spirit's internal communications capture precisely this procedural detail.
The other bidder, Mercor, illustrates even more clearly where this industry is heading.
Founded in 2023 by high school debate teammates Brendan Foody, Adarsh Hiremath, and Surya Midha, Mercor started as an AI-powered hiring platform, using algorithms to screen and interview job applicants. By September 2024, Mercor had evaluated approximately 300,000 job seekers and attained a $250 million valuation.
Soon after, the company pivoted heavily toward AI training data. By November 2025, with all three founders just 22 years old, Mercor's surging valuation made them among the world's youngest self-made billionaires. In the first half of 2026, Mercor generated over $614 million in revenue, with around 90% coming from top AI labs like OpenAI. By July, seeking a valuation near $20 billion, Mercor acquired Deeptune, a firm specializing in building simulated training environments for AI agents.
Mercor acquires training data primarily through two avenues.
First, by purchasing internal enterprise archives directly. It has approached acquired or shuttered startups to buy employee chat logs and emails, offering up to $300,000 per company.

Second, by hiring experienced professionals. TechCrunch reported in October 2025 that AI labs use Mercor to hire former corporate employees, paying them roughly $200 per hour to translate their domain expertise into training tasks and feedback. Mercor's CEO once noted that the company distributes over $1.5 million daily to individuals participating in AI training.
On one hand, it purchases corporate work logs; on the other, it hires domain experience directly from workers. Mercor's core product, at its essence, is real-world workplace execution.
This explains why Mercor showed up at Spirit's bankruptcy auction. For Mercor, 600 million internal corporate messages represented a massive reservoir of training material.
This business is not without friction. In 2025, Scale AI sued Mercor, alleging former employees had misappropriated trade secrets; subsequently, Mercor faced leaks of training data and temporary partnership freezes. Data is both Mercor's primary product and its most vulnerable asset.
A striking contrast remains: Spirit operated under its brand for 34 years; the three founders of Mercor bidding on its digital remains were just 22 years old.
The Broken Bench
Spirit's sale has not yet closed. The judge has not signed off, and Google has not received the data. Ahead lie union objections, hearings, and protracted procedural steps.
Yet this auction has brought a previously abstract issue into sharp relief.

Historically, when a company went under, much of its institutional knowledge simply vanished. Balance sheets survived, trademarks survived, patents survived. But the everyday, operational wisdom of the workforce rarely outlived the entity. How a department conducted meetings, how a manager exercised discretion, how dozens of staff coordinated during a flight delay, why a workflow evolved into its present form—these things were rarely formally preserved.
When a company disbanded, its people dispersed, inboxes were shut down, chat channels were archived, and that institutional knowledge dissolved.
In this sense, corporations have always been peculiar institutions.
A firm could operate for decades and aggregate the hard-won experience of tens of thousands of workers, yet only a minuscule fraction was ever passed down. The next company had to hire anew, make the same mistakes, and re-learn every lesson from scratch.
AI may change that permanently. If emails, meetings, support tickets, code revisions, and internal debates can be transformed into training datasets, then institutional knowledge—once locked in human heads—can finally be preserved through machine intelligence.
Where this trend will ultimately lead is difficult to foresee. Perhaps future companies will proactively archive these communications; perhaps employment contracts will be rewritten; perhaps bankruptcy statutes will introduce new privacy boundaries; perhaps corporate working records will be formally appraised in future venture rounds and M&A deals.
Spirit has brought these questions to the forefront ahead of schedule.
There is a widely shared etymological tale about the word bankruptcy. In medieval Italy, merchants conducted commerce behind wooden benches. When a trader became insolvent, his bench was broken in public. Banca rotta—the broken bench.
For centuries, when the bench broke, the story was over.
Spirit's bench is broken. But its 600 million messages are being repriced.
At a penny and a half each.