Research · Design · Develop · Implement

We build software systems that solve business problems.

Most enterprise AI stalls at the pilot. It performs well in the deck, then meets data that arrives late, contradicts itself, and was typed by a human — and quietly stops being mentioned in the steering meeting.

We build the systems that get past that point and keep running. Data engineering, advanced analytics, model design, and production AI architecture, for organizations where the answer has to hold up.

$10M
Annual saving, IBM / AK Steel expert system
5
Enterprise systems delivered into production
3
Sectors of deep domain practice
Problem → system → outcome 1 / 3
Financial
The business problem
The system we build
Outcome

Delivered results

Systems that went into production, ran against real operations, and changed a number somebody was accountable for.

$10M saved annually IBM / AK Steel

An AI expert system that emulated the judgement of a 30-year veteran operator, holding the line on decisions that had never been written down.

Behavioral fraud detection USAA

Applied AI to fraud detection, scoring customer and transaction behavior rather than relying on rules somebody had to think of first.

Payroll processing rebuilt Alight Solutions

AI-driven efficiencies in a high-volume payroll operation, changing how the processing itself worked rather than tuning around it.

Enterprise-wide cybersecurity Charles Schwab

Machine learning applied across cybersecurity and infrastructure efficiency, operating across the full production environment rather than a pilot slice.

Enterprise spend management CXO Nexus

Built AI-driven enterprise spend management systems — the platform we now implement for clients.

Managed IT & security IT Management Services

Service desk, identity and access management, endpoint security, cloud and edge platforms, and 24/7 SOC/NOC monitoring — delivered with our partner BastionX, a next-generation MSP/MSSP that treats security as the foundation and runs operations with AI-driven automation.

Why organizations partner with us

Minimized risk of failure

We work to identify tangible, measurable value early — before significant investment is committed to an approach that cannot pay for itself.

Results you can see in flight

Measurable progress at every stage rather than a single reveal at the end, so direction can be corrected while correcting is still cheap.

A genuine partnership

We work alongside your teams and transfer capability as we go. The goal is your people owning the system, not a permanent dependency on ours.

Capabilities

Disciplines that compound.

An AI system is only as good as the model beneath it, the pipeline feeding it, and the architecture holding it up. We practice all of it, because in production they cannot be separated — and where a service is better delivered by specialists, we deliver it through partners rather than pretend otherwise.

Core practice

Research & problem framing

Turning a vague business concern into a specific, measurable, solvable problem statement with agreed success metrics.

Software design & development

Product and platform engineering from architecture through delivery and implementation, built to be operated by your team.

Data engineering & pipelines

Ingestion, transformation, orchestration, labeling, and lineage at enterprise volume — with the observability to prove what happened to every row.

Advanced analytics

Statistical modeling, forecasting, and anomaly detection wired into operations rather than parked in a quarterly report.

Model design

Selecting, designing, training, and evaluating models against the business metric that matters — not against a leaderboard.

Production AI architecture

Training frameworks, deployment paths, monitoring, retraining loops, and the governance to keep a system trustworthy after launch.

Information architecture

Taxonomies, hierarchies, and canonical models designed to survive contact with real, inconsistent, human-entered data.

Innovation & thought leadership

Applied research, innovation labs, and executive briefings on where AI pays back — and where it convincingly does not.

Capability transfer

Building cross-functional teams and training internal staff to lead AI and machine learning work once we step back.

Sector depth

Domain knowledge is not optional. These are the industries where we have built systems that went into production and stayed there.

Financial services

Fraud, security, behavior

Detecting fraudulent activity through user behavior analysis; security enhancement; abnormal behavior analysis across accounts and transactions.

Insurance

Claims and cost

Streamlining claims processing, identifying anomalies in claims to cut cost, and reducing leakage before payment.

Manufacturing

Routing and uptime

Optimizing production scheduling and routing for maximum efficiency, reducing downtime and improving throughput.

Managed IT & security operations

Delivered
with BastionX

Riverton provides managed IT and security operations through our partner BastionX, a next-generation MSP/MSSP. We scope the service, set the service levels, and stay accountable for delivery. BastionX runs the operations and the security behind them. You hold one contract and one relationship, not a chain of vendors pointing at each other.

One managed practice from the service desk to the SOC — proactive IT operations and round-the-clock security, backed by AI-driven automation. Every capability below is a service BastionX delivers — see bastionx.io.

Run & support — managed IT

Managed IT

Service desk

Fast, intelligent support that simplifies IT operations and keeps your people productive.

Managed IT

Desktop & asset management

Remote troubleshooting plus total visibility into devices, licenses, and network health.

Managed IT

Patch & policy management

Automated updates that strengthen security without downtime, and centralized policy that enforces consistency.

Operational support

Proactive operations

Proactive system management and issue resolution across servers, endpoints, and networks — an extension of your team, not a ticket queue.

Physical support

On-site hands

On-site technical support for physical infrastructure and hardware when something needs a person in the room.

Continuous monitoring

SOC & NOC, 24/7

Round-the-clock security and network operations centers with smart, rule-based alerting that prioritizes what actually matters.

Secure & modernize

Cybersecurity

SOC as a service

Managed detection and response through a security operations center that hunts threats rather than forwarding alerts.

Identity & access

Verify every identity

Biometric, continuously validated identity with AI-driven trust signals (via iVALT), cutting reliance on passwords that get phished.

Zero trust

Verify-always networking

Segmented, verify-always network access so a foothold in one place does not become the run of the house.

Assessment

Penetration testing

Simulated attacks that find the vulnerabilities in your systems and defenses before someone else does.

Cloud

Cloud & edge platforms

Cloud migration and edge-platform modernization designed for resilience, not just lift-and-shift.

Compliance

Compliance automation

Automated evidence collection and continuous, audit-ready compliance in one platform (via EazyGRC) — less manual overhead at audit time.

We also implement spend analytics.

Riverton is an implementation partner for the CXO Nexus enterprise spend analytics platform, deploying it for clients who need spend visibility fast.

Visit CXO Nexus
Approach to solving business problems

A method designed to not fail quietly.

Nine steps, in order. Each one exists because skipping it is how AI and machine learning programs usually die — expensively, and long after the point where anyone could have corrected course. So each one says what skipping it costs.

The method

01

Clearly define the business problem

Identify a specific, measurable challenge, agree what success looks like, and tie it to a strategic objective before any modelling begins. Resist broadening the scope — a problem that grows during delivery is a problem nobody agreed to solve.

A poorly framed problem is the largest single cause of failure, and everything downstream inherits it.

02

Ensure data availability and quality

Assess the relevant data honestly, establish its quality, and confirm it actually contains the patterns the problem depends on.

No amount of modelling recovers from data that does not hold the answer.

03

Set up efficient data management

Build a robust pipeline for acquisition, labelling and updating, so the model is fed by a process rather than by heroics.

Without a pipeline the first model becomes the last one, because retraining is a project instead of a routine.

04

Set realistic accuracy and feasibility targets

Agree what good enough means, measured against real-world benchmarks rather than theoretical ceilings, and re-check at each stage that the solution stays feasible with the people, tools and budget actually available.

A project with no achievable target cannot tell success from failure — and perfection was never on the table.

05

Build a prototype

Prove the approach at small scale before committing to it. Test it against the target, and use what comes back to correct the problem definition and the data work while corrections are still cheap.

Untested assumptions get tested in production instead: at full cost, in front of everyone.

06

Establish training & deployment at scale

A scalable training framework and a deployment path that makes shipping a model a routine event rather than a project — sized for real volumes and real running cost.

A model that cannot be retrained or redeployed on demand starts decaying the day it ships.

07

Implement feedback loops

Regular retraining and fine-tuning, with monitoring that tells you when performance is drifting before your users do.

Without one, the system is accurate on the day it launched and quietly less so every day after.

08

Deliver outputs into the business

Get results in front of the people who make the decisions — in the systems they already use, in a form they can act on.

An accurate model nobody acts on has produced nothing. Delivery is part of the work, not an afterthought.

09

Develop in-house expertise

Build cross-functional teams and train internal staff to lead the work, so capability stays after the engagement ends.

Capability that leaves with the consultants was rented, not built.

What the method delivers

Clarity

A focused problem

A measurable challenge worth solving, with an achievable target agreed before investment.

Readiness

Data and process

Data that is ready and well managed, with the pipeline and the people to keep it that way.

Proof

A working prototype

Evidence at small scale that the approach delivers, before the commitment is made.

Scale

Production and feedback

Deployed into the systems that use it, monitored and improved as conditions change.

The outcome is the point: improved security, reduced cost, and measurably better operational efficiency — not a model that performs well in a notebook.

Let's begin with the problem.

Bring us a business challenge and we will build a roadmap tailored to your organization — starting with whether AI is the right answer at all.

Implementation partner — CXO Nexus spend analytics platform

Spend visibility, implemented properly.

CXO Nexus is the industry leader in line-item-level spend analytics, with nearly a decade in the field.

Riverton & Associates is an implementation partner for the platform. We deploy it, integrate it with your source systems, design the taxonomy it runs on, and make sure it produces answers your finance and procurement teams will actually defend.

$1T+
Of spend processed by the platform to date
$4B
Of new spend taken in every month
60+
Enterprise deployments across 10+ industries
Spend data · intake → review
After publish · who sees what
Platform by
CXO Nexus

The platform, its models, and its patented spend-management innovations are built and owned by CXO Nexus. Riverton's role is implementation partner — onboarding, source-system integration, taxonomy design, and the data engineering that gets your extracts into it cleanly.

What the platform does

Nine operations that take raw accounts-payable extracts and turn them into a governed, queryable view of every dollar. Select a stage to watch a single transaction line change — the record is illustrative, the transformations are real.

Resolution layer

Transaction line 01 · Cleanse 4 fields

Intelligence layer

05

Unmask the reseller

See through the VAR, distributor, or marketplace to the manufacturer whose product you actually bought.

06

Corporate parentage

Roll subsidiaries, acquisitions, and trading names up to the ultimate parent, so leverage is measured against the entity you negotiate with.

07

Rogue IT spend

Surface technology purchased on expense cards and departmental budgets, entirely outside the IT ledger.

08

Abnormal spend

Statistical outliers, duplicate payments, price variance across business units, and invoices that break their own historical pattern.

09

Natural-language agent

Ask the question the way you would ask a colleague. The agent resolves entities, applies parentage, and shows the rows behind the number.

Governed throughout

Every transformation is versioned and explainable — because a number you cannot defend is a number you cannot use.

Ask it like a colleague

The agent sits on the governed layer, not the raw extract. It knows which suppliers roll to the same parent, which are resellers, and which spikes were true-ups — and it will show the lines that prove it.

Illustrative exchange — figures are sample data.

Spend agent governed layer
Procurement lead
How much did we spend with Dell last year, including anything bought through resellers?
Agent
Direct spend with Dell entities was $4.18M. Including hardware bought through resellers, true Dell-manufactured spend was $6.42M54% higher than the direct figure alone.
Dell Marketing L.P. (direct)$3.41M
Dell EMC (direct)$0.77M
via CDW — unmasked$1.62M
via Insight Enterprises — unmasked$0.62M
Total, ultimate parent$6.42M
Procurement lead
Any of that bought outside IT?
Agent
Yes — $310K across 47 lines charged to Marketing, Facilities, and three regional cost centers. Mostly workstations and monitors on expense cards. None of it reached the IT asset register.

Bring us your messiest extract.

Send a representative sample and we will run it through resolution and show you what your spend actually looks like — before any commitment.

Alliance partner — BastionX · the next-generation MSP / MSSP

Amplify security, simplify IT.

BastionX is a New York–based managed IT and cybersecurity provider that treats security as foundational rather than bolted on — 24/7 monitored operations, managed defense, and AI-driven automation. Riverton partners with BastionX so the infrastructure under your data is as well-run as the analytics on top of it.

BastionX — Amplify Security, Simplify IT

Autonomous. Secure. Open. — BastionX runs an integrated SOC/NOC with U.S.-based support, turning complexity into confidence for compliance-heavy organizations. BastionX is an independent managed-service and security provider. Riverton introduces and partners with them to deliver IT operations and cybersecurity alongside our spend-analytics work.

What BastionX brings

Managed security, managed infrastructure, and the automation that makes a small, senior team behave like a much larger one — summarized from bastionx.io.

Managed security — MSSP

01

Managed SOC

A 24/7 security operations center with proactive threat hunting, detection, and incident response — not just alerts forwarded to your inbox.

02

Penetration testing

Adversarial testing that validates your defenses before someone else does, with remediation guidance your team can act on.

03

Identity & access

Passwordless, least-privilege identity and access management — closing the credential gaps that most breaches walk through.

Managed IT operations — MSP

04

Managed IT & NOC

AI-driven 24/7 network monitoring plus responsive, U.S.-based support that keeps day-to-day infrastructure quietly working.

05

Cloud & edge

Cloud migrations and edge platforms designed for resilience — moving workloads without moving the risk onto you.

06

Zero Trust networking

Segmented, verify-always network access so a foothold in one place does not become the run of the house.

Build & govern

07

DevSecOps

Security built into the software delivery pipeline — vulnerabilities caught in the build, not in production.

08

SOC 2 & compliance

Automated evidence collection and continuous compliance for SOC 2, HIPAA, and the frameworks your industry answers to.

09

Agentic AI automation

Autonomous operations that scale a senior team — routine work handled by agents so people focus on the exceptions.

Industries served

Healthcare · HIPAA Financial services Legal & professional Nonprofits Retail Energy & utilities Public sector & federal

Need the IT and security under your data handled?

BastionX runs the infrastructure and defense; Riverton makes your spend and data answerable. Together that is one accountable stack.

bastionx.io
Alliance partner — Secure Tech Solutions · cloud, data center & cybersecurity

One partner for every layer of the stack.

Secure Tech Solutions delivers full-spectrum cloud computing and data center solutions — SaaS, UCaaS, DaaS, cybersecurity, physical security, backup and disaster recovery, AI-driven custom software development, and Microsoft and multi-vendor licensing — backed by IT professional services and complete MSP/CSP account management. They serve B2B and federal clients. Riverton partners with them so the platforms your data depends on are built, secured and supported by one accountable team.

Secure Tech Solutions — Complete Cloud IT Solutions

A single trusted partner for every layer of the technology stack — cloud infrastructure, data center, security and custom software, under one account team rather than a support queue. Secure Tech Solutions Inc is an independent IT consultancy and solutions provider. Riverton introduces and partners with them to deliver cloud, infrastructure, security and applied AI work alongside our spend-analytics practice.

Core competencies

Cloud, data center, security and custom software under one account team.

Cloud & hosting

01

SaaS procurement

Software-as-a-Service procurement and cloud application licensing, managed as a portfolio rather than a scatter of individual subscriptions.

02

IaaS & PaaS

Infrastructure and Platform-as-a-Service with managed secure hosting — the platform chosen to fit the workload.

03

Data center

Colocation, managed infrastructure and secure facility hosting for the workloads that should not sit in an office cupboard.

Security & continuity

04

Cybersecurity services

Including Highly Adaptive Cybersecurity Services (HACS) and support for CMMC Level 2 and Level 3 — built for organizations that have to prove their posture.

05

Backup & disaster recovery

BDR designed so an outage is an inconvenience rather than an event — the business keeps running while the fault is fixed.

06

Physical security

Access control, surveillance and facility protection, handled by the same team that secures the network rather than a separate vendor.

Workplace & applications

07

UCaaS & CCaaS

Unified communications and contact center as a service — voice, meetings and customer contact on one platform instead of three.

08

Desktop-as-a-Service

DaaS that puts a managed, consistent desktop wherever people work, without shipping hardware to follow them.

09

Custom AI & software

AI assistants, workflow automation and application layers built into the systems a business already runs.

Licensing & managed service

10

Software licensing

Microsoft and multi-vendor licensing handled on the CSP model, with the paperwork and renewals owned rather than left to you.

11

MSP / CSP management

Complete account management across the estate — one team accountable for what is deployed, licensed and supported.

12

IT professional services

The engineering hours behind the managed service, for the projects that sit outside business-as-usual support.

What sets them apart

A team, not a queue

A single point of contact backed by a dedicated team of engineers — not a ticket routed to whoever is free.

Full-spectrum coverage

From cloud infrastructure through custom software to compliance consulting, so the seams between vendors stop being your problem.

White-glove licensing

CSP-model licensing management rather than going direct to the manufacturer — someone owns the renewals, the true-ups and the entitlement questions.

Federal-ready

Service lines aligned to NAICS codes spanning cloud, cybersecurity, telecom and consulting.

Federal readiness — NAICS codes

Cloud & software

513210 Software Publishers (SaaS)
518210 Computing Infrastructure, Data Processing & Hosting (IaaS/PaaS)
423430 Computer & Software Merchant Wholesalers

Cybersecurity & IT services

541519 Other Computer Related Services (Cybersecurity/HACS)
541513 Computer Facilities Management (MSSP/SOC)
541512 Computer Systems Design Services
541511 Custom Computer Programming (Custom AI Dev)

Telecommunications

517810 All Other Telecommunications (UCaaS/CCaaS/VoIP)
517311 Wired Telecommunications Carriers

Consulting & compliance

541611 Management & General Consulting Services
541690 Other Scientific & Technical Consulting
541211 Offices of CPAs (Compliance Audit)

At a glance

B2B & federal clients CMMC Level 2 / 3 HACS MSP / CSP model Long Island, NY In business since 1999

Need the cloud and IT underneath it built properly?

Secure Tech Solutions builds, secures and runs the platform; Riverton makes your spend and data answerable on top of it. One stack, two specialists, no gap between them.

securetsi.com
Blog

Notes from the build.

Applied research, field notes, and opinions formed by doing the work — what pays back, what quietly does not, and what we got wrong before we got it right.

Working on something similar?

If one of these is a problem you are living with, we would rather talk about yours than write about ours.

White paper

Maximizing AI and machine learning value — a practitioner's approach.

AI and machine learning return value when they are pointed at a defined problem, fed data that actually contains the answer, and built on a process someone can repeat. Most programs fail on one of those three long before a model is trained.

This paper sets out the approach: how to frame the problem, what to establish about the data before committing, the process that keeps a model useful after launch, and why the prototype comes before the roadmap.

Introduction

AI and ML capability keeps improving, and it keeps getting easier to obtain. That is not what decides the outcome. Implementing these systems takes more than adopting the technology — it takes a clear strategy for identifying the right problem, sourcing and preparing the right data, and designing something that scales and can be measured.

What follows is that work, in the order it has to happen.

1. Define a clear business problem

Every initiative starts here, and most of the ones that fail failed here. Without a well-defined challenge, even a sophisticated model produces nothing anyone can use.

Key considerations

  • Specificity. The problem must be well defined and aimed at a measurable business outcome.
  • Success metrics. Set metrics that are both relevant and attainable, so progress is visible while there is still time to act on it.
  • Avoid scope creep. Resist broadening the problem. A tight definition is what keeps the project on track.

Why it matters. A poorly defined problem produces misaligned expectations, wasted resources, and nothing of value at the end of it. A specific problem with measurable outcomes is the foundation everything else is built on.

2. Prepare the right data

A model is only as good as what trains it. Preparation is about more than cleanliness: the data has to align with the business problem and actually contain the patterns the question depends on.

Key considerations

  • Alignment with business objectives. The data must directly support the problem at hand.
  • Data quality. Clean, consistent, and representative of the domain.
  • Relevant patterns. Models predict from patterns. If the pattern is not in the data, no amount of modelling will put it there.

Why it matters. Without relevant, high-quality data the initiative cannot succeed, whatever is spent on it. This step establishes whether the right data is in place before the commitment is made.

3. Develop an effective process

Consistent results come from a process, not from effort. Data preparation, model training and scaling all have to be repeatable — and the process needs a feedback loop, so the model keeps improving instead of quietly decaying.

Key considerations

  • End-to-end process. A clear route from data preparation through training to iterative feedback.
  • Resource allocation. Dedicated teams and tools, allocated deliberately rather than borrowed.
  • Monitoring and adjustment. Watch performance, and change the process when the numbers say to.

Why it matters. A structured process is what makes output reliable rather than occasional, and it leaves room to refine and scale without starting again.

4. Build a prototype

Prototype before committing. A prototype tests the assumptions made when the problem was framed and the data assessed — at the point where being wrong is still cheap.

Key considerations

  • Proof of concept. Show that the model delivers the expected result at small scale.
  • Performance testing. Measure it; do not infer it.
  • Feedback collection. Use what the prototype returns to correct the solution before it is scaled.

Why it matters. Prototyping is how risk comes out of a programme. It is also the earliest point at which anyone outside the team can react to something real.

Key steps for maximizing AI/ML value

01

Clearly defined business problem

Invest the time upfront. Poorly framed problems are a leading cause of failure, and they surface as unmet expectations long after the budget is committed.

02

Data acquisition

Gather and validate the data that supports the problem. It has to be relevant, reliable and representative of the domain — established before the work starts, not discovered during it.

03

Solution design

Design toward the end goal: scalable, measurable, and matched to the resources available. Build for something that will change, and cut waste during development.

04

Success metrics

Set success criteria at every stage, not only at the end. Measurable goals are what let you track progress and manage expectations while the work is still in flight.

05

Accuracy and feasibility

Set realistic accuracy expectations: perfection is not achievable. Measure against real-world benchmarks, and prefer a feasible solution to a theoretically better one.

06

Feasibility and resource planning

Re-assess feasibility at each stage and plan resources accordingly. The project succeeds or fails on having the right people, tools and budget at the right time.

07

Scalability

Make sure the solution handles the volume and complexity of real operations. Evaluate input types, processing needs and running cost before scale becomes the problem.

08

Output delivery

Outputs have to reach the business and be integrated into how it works. Plan how results are presented, so the people who make decisions can act on them.

09

Maintenance and incremental improvement

These systems need maintenance. Design continuous feedback in from the start, so the solution improves on new data instead of drifting away from it.

10

Prototype and implementation

Develop incrementally: prototype first, then build toward the full solution. Iteration is what refines the result and keeps risk contained.

Conclusion

Getting value out of AI and machine learning takes planning, clear objectives, and a structured route from problem to production. The sequence is the point: define the problem, establish the data, build the process, prove it small, then scale.

None of this is exotic. It is the unglamorous work that decides whether an investment returns anything — and it is the work most often skipped.

Riverton & Associates researches, designs, develops and implements production software and AI systems, using the approach set out here.

Riverton & Associates, Inc. · Austin, Texas · riverton-assoc.com

This is the method we work by.

The Approach page sets out how these principles run as a delivery sequence on real engagements.

White paper

The vendor name problem.

One supplier, forty-three spellings, and a finance team certain they know their top ten. Why entity resolution is the least glamorous and most valuable thing in the pipeline.

Executive summary

Every spend analysis rests on a question that sounds trivial: who did we pay? In practice the answer is scattered across dozens of spellings of the same name, in several systems, in records typed by people who had no reason to care about consistency.

Until those records are resolved to one supplier, every number downstream is wrong. Category totals are wrong. Negotiating leverage is understated. Duplicate payments hide. The top-ten supplier list — the one report every finance team trusts — is confidently, quietly incorrect.

Entity resolution is the work of turning many records into one supplier. It has no demonstration value, it produces no chart anyone wants to look at, and it is the single highest-return step in a spend pipeline.

The forty-three spellings

One supplier generates variation from a dozen independent sources, and they compound.

  • Legal form. Inc, Inc., Incorporated, LLC, L.L.C., Ltd, GmbH, Pty Ltd — with the suffix, without it, and punctuated three ways.
  • Abbreviation. MKTG and Marketing. INTL, MFG, SVCS. Whoever set up the vendor master had a forty-character field and used it.
  • Case, punctuation and damage. Trailing whitespace, control characters that survived an export, encoding damage where an accent met the wrong codepage.
  • Legal entity versus trading name. The invoice says one thing, the contract another, and the cheque is made out to a third.
  • Corporate structure. Dell Marketing L.P., Dell Products L.P., EMC Corporation, VMware — four legal entities, one ultimate parent, and no string similarity between the last two and the first two.
  • Card and expense channels. Purchasing-card lines arrive as merchant descriptors: DELL 800-999-3355 TX. Same supplier, not one shared field with the AP record.
  • Multiple systems. Two ERPs from an acquisition, a procurement tool, an expense system — each with its own vendor master, IDs, and duplicates within it.
  • Time. Companies are acquired, renamed and divested. A record correct in 2019 describes an entity that no longer exists.

Forty-three spellings is not an exaggeration. It is what you get when a dozen sources of variation multiply across a few hundred thousand transaction lines.

Why the top ten is wrong

The failure is not that the numbers are slightly off. It is that they are wrong in a direction nobody checks. Fragmentation always understates a supplier: spend that belongs to one relationship is divided among its variants, and each variant ranks lower than the whole.

As recordedSpendRank
DELL MKTG L.P.$1.6M7
Dell Marketing LP$1.1M12
DELL PRODUCTS L.P.$0.8M19
Dell Mktg L.P. (duplicate vendor ID)$0.4M31
card and expense lines$0.2M
Resolved$4.1M2

Illustrative. Every row is a real supplier with real invoices.

Nothing in the unresolved view suggests an error. The report is internally consistent, reconciles to the ledger, and is wrong about the thing it exists to tell you. The consequences are specific:

  • Leverage is left on the table. You negotiate as a $1.6M customer when you are a $4.1M customer. The supplier's account team knows the real number. You do not.
  • Tail spend looks like tail spend. Fragments fall below the threshold where anyone looks, so a major relationship is managed as thirty small ones.
  • Duplicate payments survive. Detection keys on supplier plus invoice number. Two spellings are two suppliers, so the duplicate never surfaces.
  • Contract coverage is overstated. Spend under contract looks compliant because the off-contract variants count as different suppliers.
  • Risk and concentration are invisible. Concentration limits, sanctions screening and supplier risk scoring all assume you know the counterparty.

Why this is harder than string matching

The instinct is to reach for fuzzy matching. It is not sufficient, and it is not even the main event.

High similarity, different entities. Dell Marketing L.P. and Dell Products L.P. are almost identical strings and are genuinely different legal entities with different contracts and payment terms. Whether they should be merged depends on the question being asked — a business decision, not a distance threshold.

Zero similarity, same entity. IBM and International Business Machines Corporation. Alphabet and Google. Meta and Facebook. No string metric will connect them. Only reference data will.

Name similarity is therefore one signal among several, and the others carry more weight than people expect: tax identifiers, registered address, bank remittance details, the GL codes a supplier is normally posted to, the buying entity, what was bought, and the behavioural fingerprint of invoice cadence and amount distribution. A match supported by four weak signals is stronger than one supported by a single strong name score.

Scale. Naïve comparison is quadratic — a few hundred thousand vendor records is tens of billions of pairs. The work has to be blocked, and blocking is where most of the recall is won or lost: a pair that never enters a block can never be matched, no matter how good the scorer is.

The grain problem

There is no single correct answer to who is this supplier, because the right grain depends on the question.

QuestionRight grain
What leverage do we have in a negotiation?Ultimate parent
Who do we owe, and on what terms?Legal entity
Who actually delivers, and what is our risk?Operating entity or site
Which contract governs this line?Contracting entity

A layer that flattens everything to the ultimate parent cannot answer the second and third questions. One that stops at the legal entity cannot answer the first. The resolution is to keep the hierarchy rather than choose a level: resolve each record to a canonical legal entity with a stable identifier, then attach parentage above it, so reports roll up to whichever level the question needs. Choosing a single grain early is the most common irreversible mistake in this work.

The asymmetry that should govern the design

False merges and false splits are not equally bad.

A false split leaves two records unmerged. The number is understated, the fragment is visible, and someone eventually notices a familiar name in the wrong place. It is a known unknown, and it is recoverable.

A false merge combines two genuinely different suppliers. The number is now confidently wrong, it reconciles, nothing looks unusual, and it has corrupted a figure someone is about to act on. It is very hard to detect after the fact, and it destroys trust in the whole layer when it is found.

This asymmetry should shape every threshold. Bias toward precision. Route the uncertain band to human review rather than guessing. And never let an automated merge be unexplainable: when a business owner challenges a number — and they will, usually the number that matters most — the answer has to be a specific reason, not a shrug.

How it is actually done

  • Normalise deterministically first. Case, whitespace, control characters, encoding repair, legal-suffix handling, an abbreviation dictionary. Dull, rule-based and testable — and it removes a large share of the variation before any model sees the data. Probabilistic matching on unclean input wastes its discriminating power on noise.
  • Block. Partition into candidate groups on cheap keys — name prefixes, phonetic keys, tax ID, postcode, domain. Use several passes with different keys, because any single key has a blind spot.
  • Score on multiple signals. Combine them into a calibrated score, not an arbitrary weighted sum. A score that means something as a probability is what lets you set thresholds honestly.
  • Resolve to a canonical entity with a stable identifier that survives re-runs. A report whose supplier IDs change every month cannot be compared to itself.
  • Attach the hierarchy — parentage, ultimate parent, firmographics — so the grain question is answered by rolling up rather than re-resolving.
  • Queue the uncertain band for review. Make it small enough to actually clear, and capture every decision.
  • Persist decisions and feed them back. Human adjudications are the most valuable training data available. A layer that discards them re-asks the same questions every month and never improves.

How to tell whether it worked

Resist reporting percentage matched. It rewards over-merging, which is exactly the failure mode with the worst consequences. Better measures:

  • Spend under management — the share of total spend attached to a resolved supplier above the review threshold. This is the number that ties to value.
  • Precision on a sampled gold set — stratified sample, adjudicated by hand, reported honestly with the sample size.
  • Distinct supplier count before and after — a sanity signal, not a target. A large drop is expected; an implausibly large one means over-merging.
  • Stability across runs — the share of records whose supplier ID is unchanged month on month. Instability destroys trend reporting even when each run is accurate.
  • Review queue burn-down — if the uncertain band never clears, the thresholds are wrong.

Set these criteria before building, not after. A resolution layer with no agreed measure of success cannot be said to have succeeded.

Conclusion

Entity resolution produces nothing anyone wants to demonstrate. There is no visualisation, no model architecture worth discussing, and no moment where it looks impressive. It is normalisation rules, blocking keys, calibrated thresholds and a review queue.

It is also the step on which everything else depends. Classification applied to fragmented suppliers classifies fragments. Anomaly detection on fragmented suppliers finds nothing, because the baseline is split across the variants. Every downstream number inherits the quality of this one.

The finance team is not wrong to trust their top ten. They are wrong about which suppliers are in it — and they have no way to know that from the report itself. Resolving the vendor name is what turns spend data into an answer somebody can defend.

Riverton & Associates implements the CXO Nexus enterprise spend analytics platform, including the resolution layer described here.

Riverton & Associates, Inc. · Austin, Texas · riverton-assoc.com

This is the resolution layer.

The CXO Nexus platform page shows where entity resolution sits in the pipeline, and what the record looks like on either side of it.

White paper

Expert systems didn't go away.

An old idea with relevance. Before the current wave, we captured how a thirty-year veteran made a judgement call and ran it at scale. That pattern is still the highest-return AI most operations can deploy.

Executive summary

Around twenty years ago we built an expert system with IBM for AK Steel. It encoded the judgement of a thirty-year veteran operator — decisions that had never been written down — and applied them on every shift rather than on the shifts he happened to work. It saved roughly $10 million a year.

It is still running. It is still saving money.

That fact deserves more attention than it usually gets. In an industry where a production model is doing well to survive eighteen months before drift forces a retrain, a system built two decades ago is still making the same decisions correctly. This paper argues that the pattern behind it — elicit what a skilled practitioner knows, encode it, run it everywhere — remains the highest-return application of AI available to most operations, and explains where it wins, where it loses, and why the historic barrier to building one has largely fallen away.

What the pattern actually is

An expert system captures how a person makes a decision and executes that decision on every case, consistently, without fatigue.

It is worth being precise about the distinction from machine learning, because the two get confused. A learned model infers rules from labelled examples. An expert system takes the rules from a person who already has them. There is no training set, no labelling exercise, and no requirement that the past contains enough examples of the situation you care about.

That difference matters most where it is least convenient: rare events, new processes, and decisions that were never logged in a form a model could learn from.

The steel case

A thirty-year veteran operator on a production line accumulates a body of knowledge that is almost entirely undocumented. He knows which combinations of conditions precede a problem. He knows when the specification is a hard limit and when it is a guideline. He knows which readings to trust and which sensor drifts. Most of this he cannot readily articulate — asked why he made a call, the honest answer is often because that is what it needed.

The value of encoding that was not that the system out-thought him. It was that it held the line when he was not there.

An operation does not run at the level of its best operator. It runs at the average of everyone on every shift, including the new hire on nights and the stand-in during holidays. The gap between the best operator's judgement and the median decision made across a year is enormous, and it shows up as scrap, rework, downtime and lost throughput. Closing that gap — not exceeding the expert, just reaching him consistently — was worth about $10 million a year.

Why it lasted twenty years

The longevity is the part that should change how people think about this.

  • The knowledge does not drift. The system encodes process behaviour and material physics. The relationship between conditions and outcome is the same today as it was then. Contrast this with a model trained on consumer behaviour or fraud patterns, where the world moves underneath the model and yesterday's accuracy is no guarantee of today's.
  • It is inspectable. Twenty years of engineers have been able to open it, read what it does, and understand why. A rule states its own reasoning. A change can be reviewed before it ships. Nobody has ever had to take its output on trust.
  • It fails loudly. Given input outside anything anticipated, a rule simply does not fire and the system says so. It does not produce a confident, plausible, wrong answer — the characteristic failure mode of the systems currently receiving all the attention, and a genuinely dangerous property in an operation where someone acts on the output.
  • It has no infrastructure treadmill. No retraining schedule, no feature store to maintain, no accelerator to budget for, no dependency that goes end-of-life every eighteen months.

Why the pattern was abandoned anyway

It is worth being honest about this rather than nostalgic. Expert systems fell out of use for real reasons, not only fashionable ones.

  • The knowledge acquisition bottleneck. Getting rules out of an expert's head was slow and expensive — a skilled interviewer, many sessions, and an expert willing to spend weeks on it. Worse, experts genuinely cannot articulate much of what they know, so the elicitation had to infer the rule from watching decisions rather than from asking.
  • Brittleness at the edges. Systems handled the anticipated cases well and the unanticipated ones not at all.
  • Maintenance load. Rule sets grow, interact, and eventually nobody is confident about what changing one will do.
  • Over-promise. The 1980s commercial wave claimed far more than it delivered, the funding collapsed, and expert system became a phrase that made you sound twenty years out of date.

Then machine learning arrived, then deep learning, then language models, and each absorbed the attention and the budget. The pattern was not displaced because something better was found for the same problem. It was displaced because the industry's attention moved — and a technique that had stopped being interesting stopped being considered, while the systems already built carried on quietly working.

Where it still beats a learned model

  • No labelled history exists. A new line, a rare failure mode, a decision nobody recorded. A model cannot learn what the data does not contain; a person can still tell you.
  • The decision has to be explained. Regulated processes, safety cases, anything a customer or auditor will challenge. "The model scored it 0.83" is not an explanation. A fired rule is.
  • Errors are expensive and asymmetric. When being confidently wrong costs far more than declining to answer, a system that refuses to fire outside its competence is worth more than one that always produces something.
  • The input is already structured. Sensor readings, ERP fields, process telemetry. Most of the machinery of modern AI exists to cope with unstructured input; if yours is structured, you are paying for capability you do not need.
  • The expert exists and is available. Which brings us to the reason this is urgent.

Where it loses

Fairness requires stating the other half. Expert systems are the wrong choice for perception — images, audio, free text — and for high-dimensional pattern recognition where the signal genuinely is a statistical regularity nobody could articulate. They are wrong where the domain really does drift, such as fraud and pricing, and they are wrong when nobody actually knows the rules. If your best practitioner cannot outperform chance, there is nothing to elicit and you need a model that learns.

The failure to avoid is treating this as an ideological choice. It is a question about the problem: does the knowledge exist in a person, and is the domain stable? If yes, encode it. If no, learn it.

The bottleneck has moved

Here is what has actually changed, and why the pattern deserves reconsideration now rather than merely respect.

The historic cost of an expert system was elicitation — the weeks of interviews, the analyst translating them into rules, the slow round trips to validate. That cost has fallen sharply. A language model can conduct the structured interview, work through decision logs and shift notes to propose candidate rules, spot contradictions between what an expert says and what the records show, draft the first implementation, and generate the test cases that prove it.

That suggests a specific architecture, and it is the one we would build today:

  • A language model at the edges — reading unstructured input, eliciting and maintaining the rules, explaining outcomes in plain language.
  • A deterministic engine at the core — making the decision, recording which rule fired and on what evidence.

The model handles what it is genuinely good at: language, ambiguity, and interpretation. The decision itself stays in a component that is inspectable, testable, and cannot hallucinate. This gets the elicitation speed of the current wave with the durability and auditability of the old pattern — and it is a far better answer for an operations decision than putting a language model in the decision path and hoping.

Why this is the highest-return AI most operations can deploy

Three reasons, and none of them are technical.

  • The gap is consistency, not capability. Most operations already contain someone who makes the right call. The loss is not that nobody knows — it is that the knowledge is present on some shifts and absent on others. Capturing the best judgement and applying it everywhere is a large, immediate gain that requires no new science.
  • The expertise is leaving. The thirty-year veteran is, by definition, near the end of a thirty-year career. When they retire the knowledge goes with them, and no amount of subsequent investment recovers it. There is a window, it is closing in most industries right now, and it does not reopen.
  • The cost profile is unusually favourable. No labelling programme, no accelerator budget, no retraining schedule. The dominant cost is the expert's time, which is finite for reasons that have nothing to do with the project.

Conclusion

Expert systems did not fail. They went out of fashion, which is a different thing, and the industry stopped distinguishing between the two.

The system we built for AK Steel has been running for two decades and is still saving money — a claim very few modern AI deployments will be able to make in 2045. It did not out-think anyone. It took what one experienced person knew, and made it available on every shift.

That is still the most valuable thing most operations can do with AI. The tools for building it have improved considerably. The people whose judgement is worth capturing are retiring. Both of those facts point the same way.

Riverton & Associates researches, designs, develops and implements production software and AI systems, including the expert system described here.

Riverton & Associates, Inc. · Austin, Texas · riverton-assoc.com

Whose judgement is about to walk out?

If there is someone in your operation everybody defers to, that is the place to start — and the window is not open indefinitely.

White paper

Rogue IT is a measurement problem.

Technology bought outside IT is not usually defiance. It is an expense line nobody categorized, in a system nobody joined to the asset register.

Executive summary

Most organisations discuss rogue IT — shadow IT, unsanctioned spend, whatever the local term is — as a behavioural problem. Somebody went around the process. The response is a policy memo, a stricter approval workflow, and a mild sense of grievance.

That framing is wrong, and it is why the problem never goes away.

Rogue IT is overwhelmingly a measurement failure. A manager needed a tool, the sanctioned route would have taken six weeks, and a corporate card took four minutes. The purchase was neither hidden nor defiant. It posted to a departmental expense code, under an approval threshold, against a merchant descriptor that matches nothing in the vendor master — and was never joined to the asset register, so no system anywhere knows the tool exists.

You cannot govern what you cannot see. Attack rogue IT first as a data problem: find it in the spend, quantify it, and only then decide what to do about it. The enforcement conversation is unproductive until the number exists.

It is not a discipline problem

Start by discarding the moral frame. In our experience the overwhelming majority of unsanctioned technology purchases share one profile: a competent person, doing their job, choosing the fastest route to a tool they needed.

If the sanctioned path takes six weeks and a credit card takes four minutes, the organisation has expressed a preference, whatever the policy says. People followed the incentive that was actually in place.

This matters practically, not just rhetorically. If you treat it as defiance, your response is enforcement — blocking, memos, tighter thresholds — which pushes the same purchases further out of view. Personal cards. Expenses coded as "training". Purchases made by a team in another country. You have not reduced the spend; you have reduced the visibility of it, which is the opposite of the goal.

Why it is invisible

The invisibility is mechanical, and each mechanism is mundane.

  • It arrives as expense, not procurement. No purchase order, so no procurement record, so none of the controls that hang off a PO.
  • It posts to the wrong place. The GL code is departmental — marketing, operations, R&D — not technology. Any report built on spend under the IT cost centres excludes it by construction.
  • The supplier does not resolve. A card line arrives as a merchant descriptor. It shares no field with a vendor master record, so it is not recognised as a technology supplier, or as the same supplier three other departments are also paying.
  • It sits under the threshold. $400 a month is invisible in a way that $4,800 a year is not. Approval limits are per-transaction; the spend is per-month.
  • It renews silently. Auto-renewal means the decision is made once and never revisited. No one re-approves it, so no one re-examines it.
  • Nothing joins it to the asset register. The decisive one. Even where the spend is visible, the tool never enters the CMDB, the software asset register, the SSO directory or the offboarding checklist. The finance system knows you pay for it. No operational system knows it exists. Nothing connects the two.

What it costs

The direct overspend is real but it is rarely the largest number.

  • Duplicate tooling. Four departments buy four different tools that do the same thing — and often several buy the same tool on separate contracts at separate prices.
  • Lost leverage. Fragmented purchases of one vendor are invisible as a relationship, so you buy at list price repeatedly instead of negotiating once. This is the same fragmentation that breaks the supplier top-ten, arriving through a different door.
  • Retail renewal terms. Nobody negotiates a renewal they do not know is coming.
  • Uncontracted data exposure. Customer or employee data in a SaaS product with no contract, no data processing agreement, and no security review — an exposure that exists whether or not anyone has noticed it.
  • Access that outlives employment. A tool outside SSO is a tool outside offboarding. The leaver keeps their login, and no process will ever catch it.
  • An asset register that is wrong. Every downstream process that trusts it — audit, licence compliance, security posture, disaster recovery scope — inherits the error.

Only the first three are money. The rest are risk, and the risk is unquantified precisely because the spend was never measured.

Why the usual responses fail

Policy and blocking. Addresses the symptom, worsens the visibility, and does nothing about the six-week provisioning time that caused it.

Network discovery. Genuinely useful, and structurally incomplete. Traffic-based discovery misses anything used on a personal device, on a home network, or through a browser session it cannot inspect — which is most modern SaaS.

Asking people. Surveys of department heads return what they remember and are willing to declare. They often do not know either: the person who bought the tool left last year, and the card is now on someone else's expense report.

Each of these starts from the tool. The money is a better starting point, because a purchase always leaves a financial trace even when it leaves no technical one.

Measure it from the spend

The reliable path runs through accounts payable and the card programme, not the network.

  • Resolve the supplier. Merchant descriptors, expense-line free text and vendor master records all have to resolve to one canonical supplier before anything can be counted. Unglamorous entity-resolution work, and nothing downstream works without it.
  • Classify at line-item level. Not at supplier level — a supplier can sell you both technology and something else. Classification has to reach the individual line to separate a software subscription from a conference booking on the same card.
  • Enrich with what the supplier actually is. A resolved supplier can be matched against reference data to establish that it is a software vendor at all, regardless of which GL code the line was posted to.
  • Join spend to the operational registers. Take everything identified as technology spend and join it to the asset register, the CMDB and the SSO directory. The gap between what is paid for and what is known about is the rogue IT number. It is a join, not a discovery tool.
  • Flag it on the record, permanently. The output should be an attribute on the transaction — this line is technology, bought outside IT, on this cost centre — so the measurement is continuous. A study is out of date the month it lands; a flag on every line is a live control.

What to do with the number

Lead with consolidation, not enforcement. The first pass usually reveals several departments paying separately for the same product. Consolidating those is a saving nobody has to be told off to achieve, and it buys the credibility for everything that follows.

Fix the risk before the commercials. Getting a discovered tool into SSO and the offboarding process is more urgent than renegotiating its price. Access outliving employment is the exposure that turns into an incident.

Fix the provisioning time. This is the actual root cause. If the sanctioned route stays six weeks, rogue IT regenerates continuously no matter how well you measure it. A fast-track path for low-value, low-risk tools removes most of the pressure.

Set thresholds that reflect annual value. Approve on the twelve-month cost, not the monthly charge, so a $400/month subscription meets the same scrutiny as a $4,800 purchase.

How to tell whether it is working

  • Rogue spend as a share of total technology spend, trended. The absolute number matters less than the direction.
  • Time to provision a sanctioned tool. The leading indicator. If this does not fall, nothing else will hold.
  • Duplicate tool count — distinct products serving the same function.
  • Contract and DPA coverage — the share of technology suppliers with both.
  • SSO and register coverage — the share of paid-for tools that operational systems know about. The join above, reported as a percentage rather than a gap.

Conclusion

Rogue IT is not a story about people breaking rules. It is a story about an expense line that no system categorised as technology, attached to a supplier no system resolved, for a tool no system recorded.

Every part of that is a measurement failure, and every part is fixable with data work rather than policy. Resolve the supplier, classify the line, join it to the asset register, and the invisible spend becomes an ordinary number on an ordinary report.

Once it is a number, it can be managed. Until then, every conversation about it is a conversation about anecdotes.

Riverton & Associates implements the CXO Nexus enterprise spend analytics platform, which surfaces exactly this: technology spend resolved, classified at line-item level, and flagged against the cost centre that bought it.

Riverton & Associates, Inc. · Austin, Texas · riverton-assoc.com

Start with the spend, not the network.

The CXO Nexus platform resolves the supplier, classifies the line, and flags technology bought outside IT against the cost centre that bought it.

White paper

Optimizing for the wrong number.

A model that wins on the metric and loses on the outcome is a common and expensive result. Choosing what to optimize is a business decision, not a technical one.

Executive summary

The most expensive failure in applied machine learning is not a model that performs badly. It is a model that performs well — against a target nobody should have chosen.

This failure is dangerous precisely because it does not look like failure. The metric improves. The validation curve behaves. The team ships on schedule and reports a win. Six months later the business number has not moved, and nobody can explain why, because every technical indicator is still green.

Choosing the objective is the highest-leverage decision in the whole project, and it is the one most often made by default — inherited from a tutorial, a benchmark, or whatever the library optimises out of the box.

The failure looks like success

Consider the shape of it. A team is asked to reduce customer churn. They build a churn classifier, tune it hard, and reach 0.91 AUC — a genuinely good model. It goes into production. Retention runs campaigns against its output. A year later, churn is unchanged.

Nothing in the modelling was wrong. The model predicts churn accurately. The problem is that accurate churn prediction and reduced churn are different objectives, and the project optimised the first while being paid for the second.

Worse, this failure resists debugging. Nobody investigates a model that is hitting its numbers. The instinct is to tune further — more features, a bigger model, better calibration — which improves the metric again and moves the outcome not at all. Teams can spend a year on that treadmill.

Where the gap comes from

Every metric is a proxy for an outcome. Proxies diverge, and they diverge in specific, recognisable ways.

  • Accuracy on imbalanced problems. A fraud model that flags nothing is 99.9% accurate. Well known, and still ships regularly.
  • Optimising globally when you act locally. AUC measures ranking across the whole distribution. If the fraud team can investigate two hundred cases a week, the only region that matters is the top two hundred. A model with better AUC and worse precision in that slice is worse for the business and better on the report.
  • Treating all errors as equal. Minimising RMSE weights a $400 error on a small account exactly like a $400 error on the account that carries the quarter. Squared error is a mathematical convenience, not a statement about your P&L.
  • The wrong horizon. Predicting next month's churn when the retention intervention takes a quarter means every correct prediction arrives too late to act on.
  • Predicting the un-actionable. A precision-tuned churn model concentrates on the customers it is most confident about — usually the ones already gone, having cancelled the auto-renewal and stopped logging in. The customers you can still save sit in the uncertain band the model was tuned to avoid.
  • Proxy metrics with a life of their own. Click-through optimised hard enough produces clickbait. Engagement optimised hard enough produces outrage. The metric goes up, the business degrades, and the model is working exactly as specified.

The cost matrix nobody wrote down

Underneath most of these is one omission: nobody stated what the errors cost. False positives and false negatives are almost never equally expensive, and the ratio between them determines the correct operating point. That ratio is a business fact. It cannot be derived from the data, and it is not the modeller's to invent.

Churn exampleCostWhat it is
Retention offer$200what a false positive costs
Value of a saved customer$3,000what a false negative costs
Ratio1 : 15

Illustrative. Ten minutes with the business owner establishes this.

At fifteen to one, you should be willing to accept a great many false positives to avoid one false negative. Yet the default instinct — and the default threshold in most libraries — is to balance the two, and a team asked to "improve precision" will tune away from exactly the customers worth saving. Establishing this ratio changes the threshold, the metric, and sometimes the entire choice of model. It is routinely skipped.

The number has to attach to a decision

A prediction that changes nothing has no value, however accurate it is. Before choosing a metric, three questions should have concrete answers:

  • What action does this change? If the answer is "it gives us visibility", there is no decision, and therefore nothing to optimise toward.
  • Who takes that action, and what is their capacity? This sets the operating point. A team that can work two hundred cases a week defines a top-200 problem, not a whole-distribution problem.
  • How long does the action take to work? This sets the prediction horizon. A forecast that arrives after the last moment it could have been acted on is a report, not a model.

If those three cannot be answered, the metric cannot be chosen — and the honest response is to stop and answer them, not to pick a default and proceed.

Goodhart's law arrives faster than it used to

When a measure becomes a target, it ceases to be a good measure. Not a new observation, but two things have changed.

First, optimisers are much stronger than they were. A modern model asked to maximise a proxy will find the gap between that proxy and the intended outcome, and it will find it efficiently. Capability makes this worse, not better — a weak model optimising the wrong objective fails harmlessly; a strong one succeeds at the wrong thing.

Second, the loop is faster. Systems retrain on data their own decisions generated, so a proxy mismatch compounds instead of staying constant.

The practical consequence: the more capable the model, the more precisely the objective needs to be stated. Vague targets were survivable when nothing could pursue them effectively.

How to choose the number

  • Start from the decision, not the data. Write the sentence: when this number arrives, [who] will do [what] differently. If it cannot be written, stop here.
  • Write the cost matrix in currency. With the business owner in the room. Both error types, in money, even approximately — an order of magnitude is enough to move the threshold correctly.
  • Fix the operating point before training. "We will act on the top 200 per week." That sentence determines which metric is meaningful.
  • Choose the metric that matches it. If you act on a top slice, measure the top slice. If errors are asymmetric, use a cost-weighted loss. If large accounts matter more, weight by value rather than by count.
  • Add a guardrail. One metric that must not get worse, chosen to catch the proxy going feral. Optimising engagement, guard on complaints. Optimising recall, guard on the review queue.
  • Agree the counterfactual before you start. What happens without the model? Without a baseline, any result can be presented as a success.
  • Measure the outcome, not only the metric. Hold out a control group and report the business number. It is the only measurement that answers the question the project was funded to answer, and the one most often missing.

Who owns the choice

This is the argument in the title, and it is worth stating plainly.

The technical team owns how to hit the target: architecture, features, training, validation, deployment. That is their expertise and it should not be second-guessed.

The business owns what the target is: which error is more expensive, what capacity exists to act, what horizon is useful, what must not get worse. These are not technical questions. They have no technically correct answer. They are statements about how the organisation makes money and what it is willing to trade.

When the objective is delegated to the technical team — usually by omission rather than decision — it defaults to whatever is conventional. And convention in machine learning is shaped by academic benchmarks, designed for comparability across papers, not for your profit and loss. Accuracy, F1 and AUC are excellent for ranking published results. None of them knows what a false negative costs you.

How to tell it has happened to you

  • Model metrics are good and the business metric is flat.
  • The output exists and nobody uses it.
  • The team is tuning the metric rather than questioning it.
  • The metric was chosen before anyone described the decision it supports.
  • Nobody can state the cost of a false positive in currency.
  • There is no control group, so no one can say what would have happened anyway.

Any two of these together are enough to justify stopping and revisiting the objective. That conversation is uncomfortable — it implies the last several months optimised the wrong thing — and it is far cheaper than another two quarters of tuning.

Conclusion

Most failed machine learning projects do not fail technically. They hit their targets. The target was wrong, the wrongness was invisible because the metric kept improving, and the mistake was made in the first week by someone choosing a default.

Choosing what to optimise is the point where the business decides what it actually wants, in enough detail that a machine can pursue it relentlessly. That is not a task to delegate to whoever happens to be writing the training loop. It belongs to the people who know what the errors cost — and ten minutes of their time, before the modelling starts, is worth more than any amount of tuning afterwards.

Riverton & Associates researches, designs, develops and implements production software and AI systems, starting with what the system is actually for.

Riverton & Associates, Inc. · Austin, Texas · riverton-assoc.com

What would this model change?

If the answer is not a sentence about who does what differently, that is the conversation to have before any modelling starts.

White paper

Handing it over on purpose.

Designing an engagement so the client's team can run the system without us — and why that produces better architecture than building it to keep.

Executive summary

Most consultancies build to keep. Nobody says so, and it is rarely a conscious decision, but the incentives point that way and the resulting systems show it: bespoke conventions, tribal knowledge, a deploy that requires a phone call.

We design engagements the other way — so that the client's own team can run the system without us. That sounds like a commercial concession, and it partly is. But the substantive argument is different: designing for handover produces better architecture than designing to keep. Not as a happy side effect. As a direct consequence of the constraint.

Every requirement handover imposes — write it down, automate it, choose the boring option, make it observable — is something a well-built system should have anyway. Handover makes them non-optional, and supplies a test that cannot be argued with: someone who was not there has to be able to run it.

The default is retention, and it shows in the code

Building to keep does not usually look like sabotage. It looks like a series of individually reasonable decisions taken by people with no reason to decide otherwise:

  • A framework written for this project, elegant and undocumented, understood by the two people who wrote it.
  • Conventions that were never written down because everyone in the room already knew them.
  • A clever solution where a boring one would have worked, because the clever one was more interesting to build.
  • An environment that reproduces reliably on one laptop.
  • A deploy with three manual steps that are obvious if you have done them before.
  • Failure modes that are well understood and never described anywhere.

None of this requires bad intent. It is what happens in the absence of a forcing function. Nobody sets out to build a system only they can run — it is the default outcome when nobody ever has to prove otherwise. The cost lands on the client eventually, and it is larger than the invoice: a system they depend on and cannot change, staffed by people they must keep hiring back.

Handover as a forcing function

Now impose one constraint: a competent engineer who was not here must be able to run this. Watch what that eliminates.

  • Knowledge in someone's head becomes inadmissible. If it is not written down or encoded in the system, it does not exist. That converts tribal knowledge into runbooks, comments and tests.
  • Clever abstractions lose their appeal. An abstraction only pays for itself if the next person understands it faster than they would have understood the thing it replaced. Handover makes you ask that honestly.
  • Manual steps become defects. A deploy someone has to be talked through is a deploy that will eventually be done wrong.
  • Unreproducible environments become blockers. If a new engineer cannot get it running on day one, that is a bug in the system, not in the engineer.
  • Silent failure becomes unacceptable. Someone unfamiliar has to diagnose this at 3am, so alerts have to say what is wrong and what to do about it.

Read that list again with the handover framing removed. Every item is simply good engineering practice. Handover does not ask for anything that was not independently correct — it removes the option of skipping it, and supplies a test that produces a clear pass or fail.

That is the whole argument. The constraint is not a tax on the architecture. It is the quality bar, made externally verifiable.

Why end-loaded handover fails

The common pattern is to build for six months and then run a two-week "knowledge transfer" at the end. It reliably fails, for reasons that are structural rather than about effort.

  • The architecture is already shaped. By month six the system reflects who built it — the conventions, the abstractions, the assumed context. Two weeks of walkthroughs cannot unwind that.
  • Walkthroughs transfer vocabulary, not capability. The receiving team learns what the components are called. They do not learn what happens when the third one fails at month-end, because they have never seen it happen.
  • They were absent for the decisions. What survives handover badly is not what the system does — that is readable — but why it does it that way, and which alternatives were rejected. A team that sees only the outcome will re-litigate every settled question the first time something breaks.
  • It is testing at the end. Every other quality property is built in continuously and verified throughout. Handover gets treated as a phase, which is why it is the one that fails.

Handover is a property of the system, not a stage of the project.

How to actually do it

  • Their engineer deploys on day one. Not at the end. The first release to any environment should be performed by someone from the client's team, with us watching. Everything that makes that impossible is a problem to fix now rather than later.
  • Documentation is validated by use, not by writing. A runbook is correct when somebody who has not seen it follows it successfully. Have them do exactly that, and fix what they trip on.
  • Give them an incident, deliberately. Mid-engagement, when something breaks, the client's engineer leads the response and we sit behind them. The single most informative test available — and far better run while we are still there.
  • Write down the decisions, not just the code. A short record of each significant choice — what we picked, what we rejected, and why — is the part that does not survive otherwise, and it is cheap to produce at the time.
  • Prefer their stack over the optimal one. A slightly worse technology their team already operates beats a better one they have never seen. Usually the largest single architectural consequence of designing for handover, and almost always the right trade.
  • Apply the two-week test. If our whole team disappeared for a fortnight, does the system keep running, and can a change still ship? Ask it monthly. The answer is the true measure of progress.

What it costs us

Honesty requires stating the commercial side. It is slower at the start — writing decisions down, automating the deploy before it is strictly needed, pairing rather than just doing it, all cost time in the early weeks.

It also forgoes lock-in. A client who can run the system without us can also choose not to call us. That is the point, and it does reduce a certain kind of recurring revenue.

What it produces instead is better work. Clients who can run what they own come back for the next problem rather than for maintenance of the last one — and the next problem is more interesting and more valuable than the maintenance. It also removes an objection at the point of sale: a buyer who fears dependency buys less, buys later, and buys with a smaller scope.

We would rather be invited back than needed.

The architecture it produces

The systems that come out of this discipline share recognisable traits:

  • Boring technology, chosen deliberately. Conventional tools, in conventional arrangements, for the specific reason that other people already know them.
  • A smaller surface area. Every moving part is something a person must learn, which is a real cost, which makes you delete things.
  • Explicit over implicit. Convention is only free when it is the receiving team's convention. Otherwise it is a secret.
  • Observable by design. Because someone unfamiliar has to diagnose it, not because a monitoring policy required it.
  • Tests as specification. The receiving team's licence to change things. Without them, handover produces a system nobody dares touch — which is not ownership, it is custody.
  • Recorded decisions. So that inherited constraints can be distinguished from arbitrary ones.

What "done" means

Not it works in production.

Done is: their team shipped a change to it, without us, and we watched.

That is the acceptance criterion, and it is worth writing into the engagement. Anything short of it means the system is still on loan, regardless of what the status report says.

Conclusion

Designing an engagement for handover is not generosity, and framing it that way undersells it. It is the most reliable way we know to find out whether what we built is actually any good.

A system that cannot be handed over is one that has not finished being understood — its complexity is still being absorbed by the people who wrote it rather than resolved in the design. Making someone else run it is how you find that out, and doing so at the end is how you find it out too late.

Build it so they can run it without you. The architecture improves because of the constraint, not despite it.

Riverton & Associates researches, designs, develops and implements production software and AI systems — and designs the engagement so your team can run them.

Riverton & Associates, Inc. · Austin, Texas · riverton-assoc.com

Could your team run it tomorrow?

If the honest answer is no, that is worth knowing now rather than at the end of the engagement.

White paper

Agents that act, not agents that summarize.

A summary is a demo. An agent earns its place when it can take an action inside a system of record and be held to the result.

Executive summary

Most systems currently described as agents retrieve information, interpret it, and produce text. A person then reads that text and does the actual work. That is a useful assistant. It is not an agent, and the distinction is not pedantic — it is the difference between a system that demonstrates well and a system that changes a number somebody is accountable for.

The bar we hold to is specific: an agent takes an action inside a system of record, and somebody can be held to the outcome. Everything difficult about building one lives on the far side of that line, and nothing on the near side prepares you for it.

This paper is about what changes when a system stops describing the world and starts altering it: authority, reversibility, partial failure, audit, and the competence to decline.

Why summaries demo so well

A summarization demo is impressive in almost exact inverse proportion to what is at stake.

Nothing happened. No state changed, so there is nothing to reconcile, nothing to roll back, nobody to notify, and no audit trail to produce. The failure mode is a paragraph somebody disagrees with. The demo is safe because it is inert.

There is a second, subtler reason these demos flatter: summarization has no ground truth. A wrong summary looks exactly like a right one. It is fluent, plausible, and nobody in the room has read the underlying two hundred pages closely enough to catch the omission. Compare that to an agent that issues a credit note for the wrong amount — that failure announces itself, to someone, in a way nobody can talk around.

This is why pilots stall. The demo cleared a bar that had been set low by the absence of consequences, and the production version has to clear a different bar entirely.

What changes the moment it acts

Everything the demo omitted arrives at once, and none of it is model work.

  • Authority. On whose behalf is this acting, under what delegated limit? An agent inherits a permission set, and if that set is "whatever the service account can do", the blast radius is the whole system.
  • Idempotency. Networks retry. Queues redeliver. Somebody clicks twice. If the action is not idempotent, the second execution creates a duplicate payment, ticket or order — and it will happen.
  • Reversibility. Can this be undone, by whom, and within what window? Actions divide sharply into reversible and not, and that division should drive how much autonomy each is given.
  • Partial failure. It completed steps one through three of five and then the API returned a 500. The world is now in a state no one designed. Multi-step actions need explicit compensation, or they need to not be multi-step.
  • Audit. Months later somebody asks why this record was changed. The answer has to be specific: which action, on what evidence, under which rule, at whose authority. "The model decided to" does not survive contact with a regulator or a customer.
  • Reconciliation. The system of record now disagrees with three other systems that were not told. Something has to close that gap.

None of these are improved by a better model. They are systems engineering, and they are where the actual work is.

What "held to the result" requires

The second half of the claim is the harder one. Accountability is what separates an agent from a suggestion engine, and it has concrete prerequisites.

  • A named owner. A person accountable for the decisions the agent makes, in the same way they would be for a team member's. If nobody owns the outcome, nobody will investigate a bad one.
  • A measurable outcome, in business terms. Not "hours saved" and not "queries answered" — the number the process exists to move. Activity metrics are how agent programmes avoid being evaluated.
  • A justification attached to every action. Not a log line saying what happened, but a record of why: the evidence considered, the rule or threshold applied, the confidence.
  • An escalation path. Somewhere for the agent to go when it should not proceed, staffed by someone who will actually look.
  • A control group. Otherwise any change in the outcome can be attributed to the agent, and it usually will be.

If those five are not in place, the system may still be useful — but nobody is being held to anything, and its value is an assertion rather than a measurement.

The competence boundary

The most important capability an acting agent has is knowing when not to act.

An agent that acts on everything is more dangerous than one that acts on sixty per cent and escalates the rest, because the sixty per cent it handles well tells you nothing about the forty per cent it should never have touched. Coverage is a poor target; coverage at an acceptable error rate is the real one.

That requires a calibrated sense of its own confidence, an explicit threshold below which it declines, and an escalation treated as a correct outcome rather than a failure. Teams that measure "percentage handled autonomously" will tune that threshold in exactly the wrong direction.

The related discipline is the cost asymmetry: how much worse is acting wrongly than not acting at all? For a reversible, low-value action the answer may be "barely". For a payment, a customer communication, or anything touching a regulated record, the answer is "enormously" — and the threshold should reflect that rather than a default.

The decision does not belong inside the model

The architecture that survives production separates interpretation from execution.

The model interprets. It reads the messy input, works out what is being asked, and proposes an action. This is what it is genuinely good at.

A deterministic layer executes. Actions are explicit, typed, individually permissioned tools — not free-form access to an API surface. Preconditions are checked in code before anything runs. Limits are enforced outside the model, where they cannot be argued with. Every execution writes its own audit record.

The model proposes; the deterministic layer disposes. This is the same division that makes expert systems durable, arriving in a modern context: the component that decides is inspectable, testable, and unable to invent a capability it does not have.

Practically, that means an agent should not have credentials. It should have a small set of narrow tools that have credentials, each of which validates its own preconditions and refuses outside them.

How to start

  • Pick one action, not a workflow. Workflows are where multi-step partial failure lives. A single, well-chosen action delivers value and teaches you what your systems actually do under automation.
  • Run it in shadow first. The agent proposes, a person executes, and you measure agreement. This gives a real precision number before anything is at risk, and surfaces the cases nobody anticipated.
  • Graduate on a bounded slice. Let it act autonomously where the action is reversible and the value is low. Keep the caps low enough that a bad week is an inconvenience.
  • Widen by evidence. Expand scope when the reversal rate justifies it, not when the roadmap says so. This is the discipline most programmes lack.
  • Keep a human where the asymmetry demands one. Not everywhere, and not nowhere — where the cost of acting wrongly is much larger than the cost of waiting.

How to tell whether it is working

  • Reversal rate. The share of autonomous actions later corrected or undone. The precision measure that matters, and the one to report.
  • Escalation rate, and its trend. Falling escalation with a stable reversal rate is genuine progress. Falling escalation with a rising reversal rate is a threshold set too loosely.
  • Time to detect a bad action. If it is measured in months, the audit trail is not doing its job.
  • The business outcome against the control group. The only measure that answers the question the project was funded to answer.
  • Not actions taken, queries answered, or hours notionally saved. Those go up whether or not anything improved.

Conclusion

A summary is a demo. It is genuinely useful, it is easy to build, and it commits to nothing — which is exactly why it clears its bar so comfortably, and why so many of them never become anything more.

An agent earns its place when it changes something real: a record altered, a payment made, a case closed, a customer told. Everything hard about that is downstream of the moment it acts — authority, idempotency, reversal, audit, and the judgement to stop. None of it is model work, and all of it is the work.

Build the second thing. Then find out whether it was right, from a number, against a control, with somebody's name on it.

Riverton & Associates researches, designs, develops and implements production software and AI systems that take actions and are measured on them.

Riverton & Associates, Inc. · Austin, Texas · riverton-assoc.com

What would yours be allowed to do?

The useful place to start is one action, one owner, and a reversal rate you are willing to publish.

About us

A practice built on finished work.

Riverton & Associates, Inc. is an Austin, Texas software and data consultancy. We research, design, develop, and implement systems that solve defined business problems — with deep expertise in data engineering, advanced analytics, model design, and architecting AI systems for production.

What we believe

AI is a strategic asset, not a project

Applied well, it unlocks value in data that was previously impractical to reach: new products, better customer service, reduced risk, and materially more efficient operations. It enables smarter decisions, made faster.

Most efforts fail for avoidable reasons

Lack of clarity about the problem, poor execution, or goals misaligned with the business. Our methodology is built specifically to remove those failure modes rather than to demonstrate technology.

Delivered work

Systems in production at organizations where the results were measured. Several of this work also resulted in granted patents, but the patents are not the point — the systems are still running.

IBM / AK Steel
AI expert system that emulated the judgement of a 30-year veteran operator, saving $10 million annually.
Alight Solutions
Rebuilt how a high-volume payroll operation processed its work, using AI to remove effort rather than to report on it.
USAA
Applied AI to fraud detection, learning each customer's normal behavior so deviation could be scored continuously instead of caught by a rule written in advance.
Charles Schwab
Machine learning across cybersecurity and infrastructure efficiency — threat signal on one side, cost and capacity on the other.
CXO Nexus
Built AI-driven enterprise spend management systems — resolving supplier identity, categorizing spend, and surfacing what the ledger hid. We are also an implementation partner for the CXO Nexus platform.

Who we work with

We work directly with CIOs, CTOs, and CEOs in financial services, insurance, and manufacturing — leaders accountable for productivity, competitive position, and the returns on a technology budget.

Holistic

End to end

From problem definition through implementation and continuous improvement, rather than a strategy deck handed to somebody else to build.

Collaborative

Alongside your teams

Integrated with your people and your strategic vision, building in-house capability as the work proceeds.

Proven

Track record

A history of delivering measurable AI and machine learning value across regulated, operationally demanding industries.