callix.

Ten years of getting data out of documents.

Working with difficult documents?

I'm Catherine Nelson, and I've been extracting data from documents since 2016. As a Principal Data Scientist at SAP Concur I developed NLP models for extracting expense information from receipts and deployed them to production, serving hundreds of thousands of customers per day. Receipts are nasty documents, full of handwriting, poor scans and difficult layouts. I worked on these models from classical ML through RNNs to Transformer models, improving accuracy at every step. I'm also the author of two O'Reilly books on machine learning and software engineering.

Since 2023 I've worked as a consultant for startups. I've developed custom evaluation pipelines for LLM applications in healthcare, and built a document pipeline for rental leases.

If you're working with difficult documents, I can help. I can tell you when you should use an LLM and when it's overkill. I can design a document labeling strategy that will give you the results you're looking for. I can build an evaluation framework that will give you confidence in the system.