Real-time AI does not always need a GPU.
How ONNX Runtime, OpenVINO, quantization, streaming, and disciplined CPU serving met a sub-four-second production SLO while saving $1,096.83 per month.
Read article →Articles on challenges, solutions, and opinions from practical AI delivery: the messy parts of shipping models, agents, fraud analytics, automation layers, and enterprise data products.
How ONNX Runtime, OpenVINO, quantization, streaming, and disciplined CPU serving met a sub-four-second production SLO while saving $1,096.83 per month.
Read article →A field note on rebuilding an automated SAS PROC NETWORK model in open source, reducing 300M+ monthly calls into a 30-35M node graph, cutting SLA in half, and giving the client finer model control.
Read article →A field note on a fraud risk model that looked broken in production, but was actually diverging because the development and production CPUs selected different numeric execution paths.
Read article →A place for notes on feature history, customer cohorts, investigator queues, suppression logic, explainability, and why the model is usually only one part of fraud delivery.
Drafting now →An opinion piece slot for writing about useful automation: compressing repetitive handoffs while keeping human control, review, and accountability clear.
Drafting now →Good posts usually start as a real constraint, a strange failure mode, or a technical decision that needed defending. Send the thread over and I can turn it into a note.