AI systems with clear scope and measurable performance

B12Y helps organisations design and build AI systems, measure their quality, and decide what to improve, replace, or retire. Based in Auckland, we work internationally on AI and complex technical systems.

Services

Assess, design, build, and evaluate systems

We work with organisations to select, build, evaluate, and operate systems that fit their requirements.

Agentic AI Design

Design and build agents that use tools and work across systems, with clear permissions and human approval points.

AI-First Transformation

Identify where AI fits into your organisation's work, then plan the changes to workflows, responsibilities, and controls.

AI Rationalisation

Independent assessment of AI systems against operating cost, review effort, reliability, risk, and the alternatives available.

AI Quality & Evaluation

Evaluation cases, regression suites, and production measurement that make system changes visible and support reliable release decisions.

Common reasons to engage

Start with a system or decision that needs to change

B12Y can enter at assessment, design, implementation, or evaluation when the next technical decision is unclear.

01

AI experiments are not reaching production.

02

Quality cannot be measured reliably.

03

An agent needs controlled access to tools and systems.

04

Operating cost or review effort no longer justifies the system.

How an engagement works

From assessment to implementation

The scope and evidence required at each stage are agreed before implementation proceeds.

01

Diagnose

Establish the current system, the constraint, and the outcome that matters.

02

Decide

Determine whether to build, change, replace, or retire the capability.

03

Deliver

Implement and measure a system the client team can operate.

Recent technical notes

Practical writing about AI systems

Decision frameworks, implementation details, and operating considerations from current areas of work.

View all articles ->
01Design

Recovering an agent after a runtime failure

A recovery design for saving useful agent state, checking unfinished work, and controlling the return to operation.

Read article
02Optimise

Choosing a quantisation for your local model

Budget runtime memory, compare candidate files, and test whether a smaller representation still supports the workload.

Read article
03Secure

Testing what an agent can reach and change

Use disposable fixtures, enforced access controls, and independent records to test an agent's capabilities.

Read article

Discuss a technical project