AI engineering for E-commerce & Retail
AI that acts on the store, within bounds you set
We build AI that acts on the storefront and in the service queue under explicit limits: refunds, price changes, promotions, and order edits that each carry a bound, a record, and an undo.
Every money-moving action is bounded, recorded, and reversible. An agent that fails the eval does not reach the storefront.
The agent is becoming a customer, and it moves money
Two things changed at once: AI agents now shop on people's behalf, and the AI you run on your own store can refund, reprice, and cancel.
Agents are becoming the customer
Shopping agents now discover, compare, and complete purchases on a shopper's behalf, and the protocols behind it arrived this year: Google's Universal Commerce Protocol, OpenAI and Stripe's Agentic Commerce Protocol, and Shopify's agent-facing checkout tools. Agentic commerce rewards retailers whose catalogue and policies a machine can read, and a large share of retailers are not there yet.
The money-moving half is the risky half
A refund, a price change, a promotion, or an order edit is a bounded action with a price attached. Each one needs a limit, a record, and a way back, because the failure mode is money leaving the business rather than a wrong sentence.
The transparency rules are already live
Under the EU AI Act, Article 50 has been enforceable since August 2026: a chatbot has to identify itself, and AI-generated content has to be labelled. Article 5 bans manipulative techniques and exploiting a shopper's vulnerability, and the FTC's reviews rule has carried civil penalties since 2024. None of it waits for the high-risk regime, which was deferred to 2027.
Most retail AI ships unfinished
In retail CX research, roughly two thirds of teams report that at least half of their AI-powered experiences need substantial revision after launch, and testing falls away as work progresses. Retailers prioritise support automation while shoppers ask for better search and recommendations, which is how a program ends up optimising the wrong surface.
Product data decides who gets recommended
An agent shortlists from structured data, so an incomplete feed is invisible rather than merely unpersuasive. Shoppers return items when the content set the wrong expectation, and onsite personalisation has hit diminishing returns while post-purchase service stays under-invested and reliably productive.
Agentic AI fits work that carries a policy and a price
An agent can answer a status question, triage a return, or prepare a price change with the policy attached. Anything that moves money stays inside a bound you wrote.
Product discovery and content
Catalogue data, attributes, and policies structured so both a shopper and an agent can find, compare, and choose correctly, with the source attached to every claim.
- Product data enrichment
- Feed readiness
- Listing content
- Policy surfacing
Content and creative production
Product copy, catalogue content, and imagery generated at scale, with disclosure attached and a review step before anything publishes.
- Product copy
- Catalogue content
- Product imagery
- Campaign creative
Customer service and post-purchase
Order status, returns triage, warranty, and refund handling, which is where the support volume concentrates and the fastest returns usually sit.
- Order status
- Returns triage
- Warranty claims
- Refund handling
Merchandising and pricing
Assortment and long-tail pricing decisions prepared by an agent with the guardrails written in, so every price change carries a bound and a record.
- Assortment planning
- Long-tail pricing
- Promotion setup
- Markdown timing
Store and operations
BOPIS coordination, inventory exceptions, and associate copilots that work from the same data the storefront does, under the same bounds.
- BOPIS coordination
- Inventory exceptions
- Shrink signals
- Associate copilots
Featured action workflow
A governed commerce action
The agent prepares the action, policy decides whether it may run, and the customer sees what happened. Every money-moving step carries a bound, a record, and a reversal window.
-
Frame the action
Name the action, the amount at stake, and who owns it. That authority boundary is enforced at runtime rather than written into a prompt.
-
Check the bound
Limit, window, eligibility, and policy version are evaluated before anything executes, so an action outside the bound cannot proceed.
-
Retrieve the evidence
The order, the customer's own message, the returns policy, and the product record are attached to the decision rather than summarised away.
-
Decide: allow, hold, or deny
Policy returns a verdict outside the model. Routine actions run, anything uncertain waits for a person, and out-of-policy requests are refused.
AllowHoldDeny -
Execute and record
The decision record holds the reason, the policy version, the evidence, and the amount, so the action is explainable to the customer and to finance.
-
Keep the way back open
Every money-moving action is reversible for a defined window, and the customer sees the same record the agent acted on.
-
Evaluate and revalidate
Behaviour is scored against your own cases before launch and monitored after, with drift reviewed when models, policies, or catalogues change.
No client results published
What the public record already shows
We are not publishing outcomes from commerce engagements. The case for governing this work is already public: most retailers have no agent-facing manifest, roughly two thirds of retail AI experiences need substantial revision after launch, and the EU AI Act's transparency duties have been enforceable since August 2026. What follows is the standard every deployment has to meet.
The figures above are published research and regulatory dates. They are not client results, and no engagement outcome is claimed here.
- Every money-moving action bounded
- Reversal window with a record
- Disclosure at every AI surface
- Evaluation before a shopper sees it
Our path to a governed storefront
We start from one action that moves money, then build the bound, the record, and the reversal around it.
- 1
Pick the action that moves money
Choose the work with a price attached: refunds, returns, price changes, promotions, order edits. These are the actions worth governing first.
- 2
Write the bounds
Set the limit, the window, and the eligibility rules per action class, and decide what a person must approve.
- 3
Ground it in product and policy data
The agent works from your catalogue, your returns policy, and the order itself, with the source attached to what it says.
- 4
Put the decision outside the model
Policy runs in a service the model cannot reason around, and it returns allow, hold, or deny before anything executes.
- 5
Build the record and the reversal
Every action stores its reason, the policy version, and the customer's own words, and stays voidable for a defined window.
- 6
Evaluate before it meets a shopper
Behaviour is scored against your own cases before launch, with revalidation when models, policies, or catalogues change.
The controls sit outside the model
Bounds, policy, disclosure, and the record are enforced by a service the model cannot reason around, and each verdict is captured before an action runs.
Review our governance approachAction bounds
Each action class carries its own ceiling: a refund cap, a discount limit, a price floor, a window. Bounded actions are the difference between an agent that helps and an agent that gives away margin.
Policy outside the model
Eligibility, limits, and window are evaluated by a service the model cannot reason around, and it returns allow, hold, or deny before execution. The verdict is recorded either way.
Disclosure at every surface
Chat assistants identify themselves, AI-generated copy and imagery is labelled, and synthetic media is declared, which is what the live transparency rules require of a storefront.
Reversal with a record
Every money-moving action stays voidable for a defined window, and the reversal is captured with the same evidence as the action itself, so the customer and the ledger agree.
Decision records
Each run stores the reason, the policy version, the evidence it used, the amount, and who approved it, which is what makes an action explainable after the fact.
Evaluation before launch
Behaviour is scored against your own cases before an agent meets a shopper, because a retail AI experience that needs substantial revision after launch has already cost conversion.
Peak behaviour
Under load the agent degrades in a defined way: queue, hand off, or stop, with cost bounds and rate limits in place, rather than inventing an answer to keep the queue moving.
Standards and regulation
The rules that shape what a commerce deployment has to produce, and where each one lands in the architecture.
- EU AI Act Article 50
- Transparency duties for AI that talks to people: chatbots disclose, AI content is labelled. Enforceable since August 2026.
- EU AI Act Article 5
- Prohibits manipulative techniques and exploiting a shopper's vulnerability, and covert pricing that hides material information.
- Agentic commerce protocols
- UCP, ACP, and agent-facing checkout tools decide whether a catalogue can be read and transacted with by an agent.
- Product data standards
- Structured attributes and complete listings, which determine whether an agent shortlists a product at all.
- FTC reviews and testimonials rule
- Fake, bought, or suppressed reviews carry civil penalties, which extends to AI-generated reviews and testimonials.
Flytebit products fit the team building your commerce stack
Use these products when review, tests, documentation, or release control constrain the engineers behind your storefront, checkout, and integrations.
Review AI-generated changes to storefront, checkout, and integration code across eight categories before merge.
Generate tests around changed functions in pricing, cart, and payment flows, then run them in your existing pipeline.
Regenerate technical and integration documentation from the repository so the written record matches the deployed stack.
Reshape requirements, review, testing, and release around AI-assisted engineering so throughput follows generation speed.
Choose the first decision the engagement must produce
The exposure is unclear.
AI feasibility study
The action set, bounds, disclosure duties, data readiness, and a Go or No-Go verdict. Assess the workflowWe know what to build.
Architecture and delivery
A production system with bounds, policy, reversal, and an operating model. Design the systemA vendor is selling us a tool.
Independent oversight
Architecture, disclosure compliance, vendor claims, and reversal design reviewed during the build. Review the active buildThe system is already live.
LLMOps and operations
Drift reviews, policy freshness, refund accuracy, cost controls, and peak behaviour. Review the operating modelTechnical context for commerce AI
Use these guides to examine the governance, evaluation, and access work behind the page.
Questions commerce and CX leaders ask
Can an agent issue refunds on its own?
Inside a bound you set, yes. The refund runs when it sits under the cap, inside the window, and inside the policy, and it stays voidable for a defined period. Above the bound it holds for a person, and out-of-policy requests are refused rather than negotiated. The customer sees the same record the agent acted on, so nobody has to take the system's word for it.
How do we get recommended by AI shopping agents?
By being readable. Agents shortlist from structured catalogue data, attributes, availability, and policy pages, so an incomplete feed is invisible rather than merely unpersuasive. We treat agent-facing data as a first-class surface alongside the storefront, and align it with the commerce protocols your platform supports.
What do we have to disclose?
Under the EU AI Act, a chat assistant has to identify itself as AI, and AI-generated content has to be labelled, with the duties enforceable since August 2026 and not deferred. In practice that means disclosure is a property of the surface rather than a line in the terms: the assistant says what it is, generated copy is marked, and synthetic imagery is declared.
Is dynamic pricing allowed?
Retail pricing is not in the high-risk tier, and that is where the reassurance ends. Article 5 prohibits manipulative techniques and covert pricing that hides material information, and personalised pricing carries its own disclosure rules under consumer law. Our position is simpler to operate: every price change carries a bound, a policy version, and a record, so it can be explained to a customer and to a regulator.
Where does the ROI actually come from?
Post-purchase service is the under-invested one. Order status questions are the bulk of support volume, returns triage and warranty automation contain most of it, and per-interaction cost is an order of magnitude below a human agent. Onsite personalisation is now table stakes with shrinking incremental returns, which is why we start with the money-moving and service work instead.
What happens when the agent gets it wrong?
The reversal window is the first answer, and the record is the second, because it shows what the agent saw and which policy version applied. The third is the eval that should have caught it: a failure that reaches a customer is a gap in the test set, and it goes back in.
How do we measure whether this works?
At the action level: containment on order-status and returns contact, refund accuracy against policy, reversal rate, contact deflection, disclosure coverage on AI surfaces, conversion and margin per order, and cost per interaction. Peak behaviour is measured separately, under load.
Can you review an AI vendor we are already evaluating?
Yes. An oversight engagement reviews the architecture, bounds, disclosure compliance, reversal design, evaluation approach, and vendor claims while the platform is being implemented, with findings delivered at agreed milestones.
Find the first action worth putting under a bound
We will examine the action, the bound, the disclosure duty, the reversal, and the failure modes before recommending a build.