Hey HN - I made this neat library to remove PII and other kinds of personal information from text without paying for a ML classifier. Sensored is mostly regex-based detection with features built-in to help steer LLMs to add stable anchors, light weight bloom filters + optional Jev classification, and more. It's also fast: ~0.0.8ms chat and ~12Mi/s for streaming.
I'm curious how this works under real loads, what gaps exist, and what false-positives you've seen working with tools like this in the past.
I'm curious how this works under real loads, what gaps exist, and what false-positives you've seen working with tools like this in the past.