Agentic AI Design
Design and build agents that use tools and work across systems, with clear permissions and human approval points.
B12Y helps organisations design and build AI systems, measure their quality, and decide what to improve, replace, or retire. Based in Auckland, we work internationally on AI and complex technical systems.
We work with organisations to select, build, evaluate, and operate systems that fit their requirements.
Design and build agents that use tools and work across systems, with clear permissions and human approval points.
Identify where AI fits into your organisation's work, then plan the changes to workflows, responsibilities, and controls.
Independent assessment of AI systems against operating cost, review effort, reliability, risk, and the alternatives available.
Evaluation cases, regression suites, and production measurement that make system changes visible and support reliable release decisions.
Common reasons to engage
B12Y can enter at assessment, design, implementation, or evaluation when the next technical decision is unclear.
AI experiments are not reaching production.
Quality cannot be measured reliably.
An agent needs controlled access to tools and systems.
Operating cost or review effort no longer justifies the system.
How an engagement works
The scope and evidence required at each stage are agreed before implementation proceeds.
Establish the current system, the constraint, and the outcome that matters.
Determine whether to build, change, replace, or retire the capability.
Implement and measure a system the client team can operate.
Recent technical notes
Decision frameworks, implementation details, and operating considerations from current areas of work.
A recovery design for saving useful agent state, checking unfinished work, and controlling the return to operation.
Read articleBudget runtime memory, compare candidate files, and test whether a smaller representation still supports the workload.
Read articleUse disposable fixtures, enforced access controls, and independent records to test an agent's capabilities.
Read article