I build the feature and the measurement that says whether it actually works: accuracy with an interval honest about the sample size, cost per correct answer, the failure cases by name, and the threshold to gate it at. Every project on my site publishes its losing axis. Yours will too, because a result you can't interrogate is marketing, not engineering.
Four shapes of work. If yours is close to one of these, it's probably a fit; if it isn't, say so on the call and I'll tell you honestly whether I'm the right person.
You have something in production and no number you trust. I build a frozen task set from your real data, hand-label it, and report accuracy next to cost and latency, with a confidence interval that admits how small the sample is. You get the failure taxonomy, the threshold that trades wrong answers against coverage, and the harness so you can re-run it after every change.
Agents, retrieval, extraction, classification, the API and UI around them. Delivered with the harness that proves it works, not a screen recording. I've shipped this end to end at a GPU cloud and at a medical imaging startup: identity resolution over evidence chains, LLM tiering with human-review gates and audit trails, and a sixteen-tool conversational agent over a live database.
Routing and cascades, batching, quantization, the boring wins that show up on the invoice. Measured against a frozen task set and ranked by dollars per solved task, so the saving is a fact rather than a hope. My verify-then-escalate cascade beat every single-model baseline on 100 EvalPlus tasks at 47.1% of the cost, and GMI Cloud published it.
Gateways, rate limiting, queues, storage, the deploy path. Go and C++ and Python, on Kubernetes and Terraform, load-tested before I hand it over rather than after you find out. If it needs to hold a global limit across replicas, or keep every acknowledged write through a power cut, that is the kind of thing I build for fun.
No discovery phase you pay for, no hourly meter running while I read your codebase.
You describe the system and what's uncertain about it. I tell you what I'd measure or build first, and whether I'm the right person. Sometimes the answer is that you need an afternoon and a spreadsheet, not me, and I'd rather say that than take the work.
What I'll deliver, the criterion that decides whether it worked, and the date. Priced per engagement, not per hour, so the incentive is to finish rather than to linger. You approve the page before anything starts.
Code in your repository, commented the way the surrounding code is commented, plus the measurement that says whether it worked, with the losing axis included. If the answer is that the feature doesn't clear the bar, you get that answer with the evidence, which is worth more than a green dashboard.
Questions, small fixes, and help re-running the harness once your data has moved. Included, not billed.
I take one engagement at a time alongside school and my internship, so I'll tell you my real availability on the first call rather than after you've committed. Email me at lakshgoyal06@gmail.com, or leave it here and I'll come back to you.