AI Can’t Find Everything Buried in Corporate Documents. This Tampa Startup Has a Fix.

You hit send, the email disappears from the screen and the exchange feels finished, but inside enterprises that message may sit for years beside invoices, contracts, benefit records, PDFs and millions of other files kept because somebody may need them again. Artificial intelligence has given companies another reason to keep all of that information within reach, but the company name on an invoice, the terms buried in a contract or the relevant paragraph inside a PDF may still exist only as rows, boxes and text rather than database fields. Before an AI model can interpret that information, somebody — or something — still has to find it, pull it out and put it into a structured form the next system can understand.

JPMorgan Chase is working through that problem across one of the world’s largest corporate data estates. The bank says it holds more than an exabyte of structured and unstructured information, voice and video, and Chief Data Officer Mark Birkhead has described making that information “AI ready” as a multiyear priority, including publishing data consistently enough for large language models to use without employees continually repairing what sits underneath them.

Advertisement

Jordan Rhoads, CEO of Tampa-based nCore Data Management, Inc., is trying to sell companies another way into that first step. His company is commercializing ANNA® — short for Adaptive Neural Network Accelerator and marketed as its flagship Smart File Management System — software that runs on CPUs and lets a nontechnical employee define the fields a company wants, train the system on its documents in one pass and extract the results in a format of choice.

“There are three levels to the world we live in: phase one, phase two, phase three. Phase one is data intake; you can’t do anything with 80% of unstructured raw data enterprises possess until it’s normalized and structured,” Rhoads said. “Phase two is business intelligence and analytics, which pulls from phase one to tell you what you have. Phase three is the AI models that interpret it. Nvidia and the companies building LLMs and GPUs own phase two and phase three — that’s where GPUs and LLMs are critical. We sit at phase one, the data-intake layer, where we make the data AI-ready before it ever gets there.”

Large language models already do some of that extraction work. Uber built a generative-AI system around GPT-4 to extract information from supplier invoices and reported 90% overall accuracy, with 35% of submitted invoices reaching 99.5% and the remaining 65% exceeding 80%. The information still passes through post-processing, validation and human review before entering Uber’s systems, while an earlier fine-tuned model sometimes hallucinated invoice line items.

Concentrix processes another 100,000 utility invoices each month. Microsoft said its earlier custom models delivered roughly 65% to 70% extraction accuracy, compared with 99% for its prompt-based system by January 2026, although invoices outside the automated process still enter an exception workflow. The deployment cut the team responsible for model pattern analysis and maintenance from roughly 40 people to 11, before accounting for the compute and token costs of running the system. Uber, valued at roughly $150 billion, and Concentrix, a multibillion-dollar business-process outsourcer, operate at a scale that can absorb that overhead; nCore is targeting midmarket companies and individual divisions inside larger ones that may not have dedicated AI engineering teams, GPU budgets or employees assigned to maintain a model once it goes live.

“Companies are pointing LLMs at this and saying, in effect, ‘You’re pretty good at a lot of things so go turn this unstructured data into structured data,'” Rhoads said. “They do an OK job; we’ve seen accuracy as high as about 96%, as low as 80% or lower. But when you’re talking about regulatory, finance, health benefits, banking, anything short of 100% is unacceptable. Getting closer to that number with GPUs and LLMs gets expensive fast, in both hardware and token costs, but LLMs are probabilistic, not deterministic. We run a CPU-native, deterministic extraction tool that a nontechnical person can use, with no GPUs and no developers, with checksum deterministic results.”

Chart showing actual and forecast global data creation from 2010 to 2035 in zettabytes
Global data creation is projected to reach 2,142 zettabytes by 2035, up from 175 zettabytes in 2025.

From Government Technology to a Commercial Company

ANNA’s underlying technology predates the current AI boom. Rhoads said it was used on government and Department of Defense projects whose constraints ranged from searching compressed medical records to squeezing more information through limited transmission systems. A Cornell and LifeNET child-abuse-prevention project reduced thousands of pages of medical information by up to 84.5%, allowing physicians to search records on handheld devices without first decompressing them, while a NASA Mars-rover project reduced small packetized data by up to roughly 80% to increase transmission bandwidth, which Rhoads said helped extend the rover’s battery life by approximately two years.

In Afghanistan, Rhoads said a biometric-verification system compressed roughly 10 times as much verification information onto a standard barcode without requiring a chip, RFID or internet connection. The technology was also deployed within FEMA and the National Guard before becoming part of the organization’s entry in the U.S. Army’s Cyber Quest 2020 call for innovation.

Rhoads had come aboard as a passive investor after working in technology, medtech and medical devices for Fortune 500 and venture-backed companies. He became increasingly involved in day-to-day operations, then watched COVID-19 shut down much of the government’s work only months after the organization won its category at Cyber Quest 2020. He decided to form a C corporation, raise capital and commercialize the technology, leaving his steady job to move forward with seven employees and the company’s investors.

“It’s been a roller coaster, and there are plenty of hard days, but every morning I say, ‘I get to do this instead of I have to do this,'” Rhoads said. “‘Have to’ turns a hard day into something you’re enduring. ‘Get to’ turns the same hard day into part of building something disruptive, something that’s gaining traction and solving real problems.”

The government work had been built around defined missions and technical users. Selling to businesses meant the software had to process ordinary corporate documents repeatedly without engineers rebuilding the workflow for every new file type, format or customer.

The First Commercial Test

At HUB International, that meant complex, disparate health-benefit invoices that employees were processing one field at a time: open the document, find the number, copy it into the appropriate output and move to the next field. Standard claims warehouses capture adjudicated health claims, but pharmacy rebates, stop-loss refunds and carrier and TPA fees can arrive separately as invoices, leaving those costs in an employer’s books as lump sums rather than itemized information that can be analyzed alongside claims data. Rhoads estimates that category at 10% and growing and said that, if a representative share of the roughly $1 trillion national self-funded health insurance market follows the same pattern, about $100 billion annually would sit outside the normal claims-data pipeline.

HUB was already tackling the invoices manually in one office before nCore built ANNA out for additional offices and, eventually, the broader market. Rhoads said ANNA structured one batch of information for employer health-plan reporting in 47 seconds with checksum verification, replacing work that had required several employees over multiple days. Those employees remained, he said, but could spend less time copying information and more time interpreting it for the brokers and producers they supported.

Rhoads cautions companies against eliminating workers after automating a process at 80% or 96% accuracy, only to bring people back later to review whatever the system missed. “Don’t get rid of those people,” he said. “They’re going to be needed for sure.” Five other entities are testing or applying ANNA, according to Rhoads.

Making the Archive Smaller Without Losing Insights

Extraction addresses what companies can pull from their files. nCore’s other technology, TextComp, addresses how much room those files occupy as enterprise data is expected to roughly triple by 2028. Rhoads points to a data-storage market pushing toward $1 trillion while data compression remains below $10 billion, and said traditional compression can create another problem: keeping compressed information searchable can require a separate index that approaches the size of the storage the compression just saved.

TextComp is designed to search qualifying compressed data without first decompressing it or maintaining that separate index. Rhoads said it produces an average reduction of around 75%, allowing a company to search the compressed archive, locate a particular file and decompress only that file if the original is needed, such as during a legal proceeding.

“With TextComp, you can compress the data and get around a 75% reduction on the average and then search the compressed data directly — no indexing, no decompressing — to extract specific information or data intelligence, maybe to locate a file, and then actually decompress that specific file if you need to,” Rhoads said. “We proved this within our government work.”

TextComp has produced reductions above 93% on some text files, including JSON files, Rhoads said, but those results do not apply to everything stored inside a corporate archive or data center. Video, audio and images behave differently and can occupy far more space, while the savings for any particular company depend on how much qualifying text it stores, how many duplicate and backup copies it keeps, what indexes it maintains and how often its systems decompress and recompress information. At scale, Rhoads said, the savings could mean fewer server purchases, fewer CPU cycles or lower storage costs.

“Everyone’s getting caught up in the boardroom over AI right now, and they should be, we need to invest in it,” Rhoads said. “But what’s getting overlooked is that the cost can exceed the benefits, directly and indirectly.”

Read More
Explore the Latest Tampa Bay Business News
Read the latest coverage of business, commercial real estate, development, technology, finance and leadership from across the Tampa Bay region.