LLM features: what we check first
Most teams worry about the model saying something embarrassing. The bigger risk is what the model can do and see for the user who is typing.
1. Tools the model can call
If the assistant can look up an order, send an email or query a database, each tool is an API endpoint. We test it like one: can a user get the tool to act on someone else's data by asking nicely?
tool: get_invoice(invoice_id) check: does it verify invoice.tenant == session.tenant?
2. What it can retrieve
Retrieval pipelines often index everything and filter later, if at all. We check whether one customer's document can show up in another customer's answer.
Fix
Filter by tenant and role at query time, inside the vector store, not after the model has read the results.
3. Where its output goes
Model output rendered as HTML, passed to a shell or used to build a query is user input in disguise. Treat it that way.
Shipping an LLM feature?Book a call