Minimized risk of failure
We work to identify tangible, measurable value early — before significant investment is committed to an approach that cannot pay for itself.
Most enterprise AI stalls at the pilot. It performs well in the deck, then meets data that arrives late, contradicts itself, and was typed by a human — and quietly stops being mentioned in the steering meeting.
We build the systems that get past that point and keep running. Data engineering, advanced analytics, model design, and production AI architecture, for organizations where the answer has to hold up.
Systems that went into production, ran against real operations, and changed a number somebody was accountable for.
An AI expert system that emulated the judgement of a 30-year veteran operator, holding the line on decisions that had never been written down.
Applied AI to fraud detection, scoring customer and transaction behavior rather than relying on rules somebody had to think of first.
AI-driven efficiencies in a high-volume payroll operation, changing how the processing itself worked rather than tuning around it.
Machine learning applied across cybersecurity and infrastructure efficiency, operating across the full production environment rather than a pilot slice.
Built AI-driven enterprise spend management systems — the platform we now implement for clients.
Service desk, identity and access management, endpoint security, cloud and edge platforms, and 24/7 SOC/NOC monitoring — delivered with our partner BastionX, a next-generation MSP/MSSP that treats security as the foundation and runs operations with AI-driven automation.
We work to identify tangible, measurable value early — before significant investment is committed to an approach that cannot pay for itself.
Measurable progress at every stage rather than a single reveal at the end, so direction can be corrected while correcting is still cheap.
We work alongside your teams and transfer capability as we go. The goal is your people owning the system, not a permanent dependency on ours.
An AI system is only as good as the model beneath it, the pipeline feeding it, and the architecture holding it up. We practice all of it, because in production they cannot be separated — and where a service is better delivered by specialists, we deliver it through partners rather than pretend otherwise.
Turning a vague business concern into a specific, measurable, solvable problem statement with agreed success metrics.
Product and platform engineering from architecture through delivery and implementation, built to be operated by your team.
Ingestion, transformation, orchestration, labeling, and lineage at enterprise volume — with the observability to prove what happened to every row.
Statistical modeling, forecasting, and anomaly detection wired into operations rather than parked in a quarterly report.
Selecting, designing, training, and evaluating models against the business metric that matters — not against a leaderboard.
Training frameworks, deployment paths, monitoring, retraining loops, and the governance to keep a system trustworthy after launch.
Taxonomies, hierarchies, and canonical models designed to survive contact with real, inconsistent, human-entered data.
Applied research, innovation labs, and executive briefings on where AI pays back — and where it convincingly does not.
Building cross-functional teams and training internal staff to lead AI and machine learning work once we step back.
Domain knowledge is not optional. These are the industries where we have built systems that went into production and stayed there.
Detecting fraudulent activity through user behavior analysis; security enhancement; abnormal behavior analysis across accounts and transactions.
Streamlining claims processing, identifying anomalies in claims to cut cost, and reducing leakage before payment.
Optimizing production scheduling and routing for maximum efficiency, reducing downtime and improving throughput.
Riverton provides managed IT and security operations through our partner BastionX, a next-generation MSP/MSSP. We scope the service, set the service levels, and stay accountable for delivery. BastionX runs the operations and the security behind them. You hold one contract and one relationship, not a chain of vendors pointing at each other.
One managed practice from the service desk to the SOC — proactive IT operations and round-the-clock security, backed by AI-driven automation. Every capability below is a service BastionX delivers — see bastionx.io.
Run & support — managed IT
Fast, intelligent support that simplifies IT operations and keeps your people productive.
Remote troubleshooting plus total visibility into devices, licenses, and network health.
Automated updates that strengthen security without downtime, and centralized policy that enforces consistency.
Proactive system management and issue resolution across servers, endpoints, and networks — an extension of your team, not a ticket queue.
On-site technical support for physical infrastructure and hardware when something needs a person in the room.
Round-the-clock security and network operations centers with smart, rule-based alerting that prioritizes what actually matters.
Secure & modernize
Managed detection and response through a security operations center that hunts threats rather than forwarding alerts.
Biometric, continuously validated identity with AI-driven trust signals (via iVALT), cutting reliance on passwords that get phished.
Segmented, verify-always network access so a foothold in one place does not become the run of the house.
Simulated attacks that find the vulnerabilities in your systems and defenses before someone else does.
Cloud migration and edge-platform modernization designed for resilience, not just lift-and-shift.
Automated evidence collection and continuous, audit-ready compliance in one platform (via EazyGRC) — less manual overhead at audit time.
Riverton is an implementation partner for the CXO Nexus enterprise spend analytics platform, deploying it for clients who need spend visibility fast.
Nine steps, in order. Each one exists because skipping it is how AI and machine learning programs usually die — expensively, and long after the point where anyone could have corrected course. So each one says what skipping it costs.
Identify a specific, measurable challenge, agree what success looks like, and tie it to a strategic objective before any modelling begins. Resist broadening the scope — a problem that grows during delivery is a problem nobody agreed to solve.
A poorly framed problem is the largest single cause of failure, and everything downstream inherits it.
Assess the relevant data honestly, establish its quality, and confirm it actually contains the patterns the problem depends on.
No amount of modelling recovers from data that does not hold the answer.
Build a robust pipeline for acquisition, labelling and updating, so the model is fed by a process rather than by heroics.
Without a pipeline the first model becomes the last one, because retraining is a project instead of a routine.
Agree what good enough means, measured against real-world benchmarks rather than theoretical ceilings, and re-check at each stage that the solution stays feasible with the people, tools and budget actually available.
A project with no achievable target cannot tell success from failure — and perfection was never on the table.
Prove the approach at small scale before committing to it. Test it against the target, and use what comes back to correct the problem definition and the data work while corrections are still cheap.
Untested assumptions get tested in production instead: at full cost, in front of everyone.
A scalable training framework and a deployment path that makes shipping a model a routine event rather than a project — sized for real volumes and real running cost.
A model that cannot be retrained or redeployed on demand starts decaying the day it ships.
Regular retraining and fine-tuning, with monitoring that tells you when performance is drifting before your users do.
Without one, the system is accurate on the day it launched and quietly less so every day after.
Get results in front of the people who make the decisions — in the systems they already use, in a form they can act on.
An accurate model nobody acts on has produced nothing. Delivery is part of the work, not an afterthought.
Build cross-functional teams and train internal staff to lead the work, so capability stays after the engagement ends.
Capability that leaves with the consultants was rented, not built.
A measurable challenge worth solving, with an achievable target agreed before investment.
Data that is ready and well managed, with the pipeline and the people to keep it that way.
Evidence at small scale that the approach delivers, before the commitment is made.
Deployed into the systems that use it, monitored and improved as conditions change.
The outcome is the point: improved security, reduced cost, and measurably better operational efficiency — not a model that performs well in a notebook.
Bring us a business challenge and we will build a roadmap tailored to your organization — starting with whether AI is the right answer at all.
CXO Nexus is the industry leader in line-item-level spend analytics, with nearly a decade in the field.
Riverton & Associates is an implementation partner for the platform. We deploy it, integrate it with your source systems, design the taxonomy it runs on, and make sure it produces answers your finance and procurement teams will actually defend.
The platform, its models, and its patented spend-management innovations are built and owned by CXO Nexus. Riverton's role is implementation partner — onboarding, source-system integration, taxonomy design, and the data engineering that gets your extracts into it cleanly.
Nine operations that take raw accounts-payable extracts and turn them into a governed, queryable view of every dollar. Select a stage to watch a single transaction line change — the record is illustrative, the transformations are real.
Resolution layer
Intelligence layer
See through the VAR, distributor, or marketplace to the manufacturer whose product you actually bought.
Roll subsidiaries, acquisitions, and trading names up to the ultimate parent, so leverage is measured against the entity you negotiate with.
Surface technology purchased on expense cards and departmental budgets, entirely outside the IT ledger.
Statistical outliers, duplicate payments, price variance across business units, and invoices that break their own historical pattern.
Ask the question the way you would ask a colleague. The agent resolves entities, applies parentage, and shows the rows behind the number.
Every transformation is versioned and explainable — because a number you cannot defend is a number you cannot use.
The agent sits on the governed layer, not the raw extract. It knows which suppliers roll to the same parent, which are resellers, and which spikes were true-ups — and it will show the lines that prove it.
Illustrative exchange — figures are sample data.
| Dell Marketing L.P. (direct) | $3.41M |
| Dell EMC (direct) | $0.77M |
| via CDW — unmasked | $1.62M |
| via Insight Enterprises — unmasked | $0.62M |
| Total, ultimate parent | $6.42M |
Send a representative sample and we will run it through resolution and show you what your spend actually looks like — before any commitment.
BastionX is a New York–based managed IT and cybersecurity provider that treats security as foundational rather than bolted on — 24/7 monitored operations, managed defense, and AI-driven automation. Riverton partners with BastionX so the infrastructure under your data is as well-run as the analytics on top of it.
Autonomous. Secure. Open. — BastionX runs an integrated SOC/NOC with U.S.-based support, turning complexity into confidence for compliance-heavy organizations. BastionX is an independent managed-service and security provider. Riverton introduces and partners with them to deliver IT operations and cybersecurity alongside our spend-analytics work.
Managed security, managed infrastructure, and the automation that makes a small, senior team behave like a much larger one — summarized from bastionx.io.
Managed security — MSSP
A 24/7 security operations center with proactive threat hunting, detection, and incident response — not just alerts forwarded to your inbox.
Adversarial testing that validates your defenses before someone else does, with remediation guidance your team can act on.
Passwordless, least-privilege identity and access management — closing the credential gaps that most breaches walk through.
Managed IT operations — MSP
AI-driven 24/7 network monitoring plus responsive, U.S.-based support that keeps day-to-day infrastructure quietly working.
Cloud migrations and edge platforms designed for resilience — moving workloads without moving the risk onto you.
Segmented, verify-always network access so a foothold in one place does not become the run of the house.
Build & govern
Security built into the software delivery pipeline — vulnerabilities caught in the build, not in production.
Automated evidence collection and continuous compliance for SOC 2, HIPAA, and the frameworks your industry answers to.
Autonomous operations that scale a senior team — routine work handled by agents so people focus on the exceptions.
Industries served
BastionX runs the infrastructure and defense; Riverton makes your spend and data answerable. Together that is one accountable stack.
Secure Tech Solutions delivers full-spectrum cloud computing and data center solutions — SaaS, UCaaS, DaaS, cybersecurity, physical security, backup and disaster recovery, AI-driven custom software development, and Microsoft and multi-vendor licensing — backed by IT professional services and complete MSP/CSP account management. They serve B2B and federal clients. Riverton partners with them so the platforms your data depends on are built, secured and supported by one accountable team.
A single trusted partner for every layer of the technology stack — cloud infrastructure, data center, security and custom software, under one account team rather than a support queue. Secure Tech Solutions Inc is an independent IT consultancy and solutions provider. Riverton introduces and partners with them to deliver cloud, infrastructure, security and applied AI work alongside our spend-analytics practice.
Cloud, data center, security and custom software under one account team.
Cloud & hosting
Software-as-a-Service procurement and cloud application licensing, managed as a portfolio rather than a scatter of individual subscriptions.
Infrastructure and Platform-as-a-Service with managed secure hosting — the platform chosen to fit the workload.
Colocation, managed infrastructure and secure facility hosting for the workloads that should not sit in an office cupboard.
Security & continuity
Including Highly Adaptive Cybersecurity Services (HACS) and support for CMMC Level 2 and Level 3 — built for organizations that have to prove their posture.
BDR designed so an outage is an inconvenience rather than an event — the business keeps running while the fault is fixed.
Access control, surveillance and facility protection, handled by the same team that secures the network rather than a separate vendor.
Workplace & applications
Unified communications and contact center as a service — voice, meetings and customer contact on one platform instead of three.
DaaS that puts a managed, consistent desktop wherever people work, without shipping hardware to follow them.
AI assistants, workflow automation and application layers built into the systems a business already runs.
Licensing & managed service
Microsoft and multi-vendor licensing handled on the CSP model, with the paperwork and renewals owned rather than left to you.
Complete account management across the estate — one team accountable for what is deployed, licensed and supported.
The engineering hours behind the managed service, for the projects that sit outside business-as-usual support.
A single point of contact backed by a dedicated team of engineers — not a ticket routed to whoever is free.
From cloud infrastructure through custom software to compliance consulting, so the seams between vendors stop being your problem.
CSP-model licensing management rather than going direct to the manufacturer — someone owns the renewals, the true-ups and the entitlement questions.
Service lines aligned to NAICS codes spanning cloud, cybersecurity, telecom and consulting.
Federal readiness — NAICS codes
513210 Software Publishers (SaaS)
518210 Computing Infrastructure, Data Processing & Hosting (IaaS/PaaS)
423430 Computer & Software Merchant Wholesalers
541519 Other Computer Related Services (Cybersecurity/HACS)
541513 Computer Facilities Management (MSSP/SOC)
541512 Computer Systems Design Services
541511 Custom Computer Programming (Custom AI Dev)
517810 All Other Telecommunications (UCaaS/CCaaS/VoIP)
517311 Wired Telecommunications Carriers
541611 Management & General Consulting Services
541690 Other Scientific & Technical Consulting
541211 Offices of CPAs (Compliance Audit)
At a glance
Secure Tech Solutions builds, secures and runs the platform; Riverton makes your spend and data answerable on top of it. One stack, two specialists, no gap between them.
Applied research, field notes, and opinions formed by doing the work — what pays back, what quietly does not, and what we got wrong before we got it right.
Nothing matches that search.
If one of these is a problem you are living with, we would rather talk about yours than write about ours.
AI and machine learning return value when they are pointed at a defined problem, fed data that actually contains the answer, and built on a process someone can repeat. Most programs fail on one of those three long before a model is trained.
This paper sets out the approach: how to frame the problem, what to establish about the data before committing, the process that keeps a model useful after launch, and why the prototype comes before the roadmap.
AI and ML capability keeps improving, and it keeps getting easier to obtain. That is not what decides the outcome. Implementing these systems takes more than adopting the technology — it takes a clear strategy for identifying the right problem, sourcing and preparing the right data, and designing something that scales and can be measured.
What follows is that work, in the order it has to happen.
Every initiative starts here, and most of the ones that fail failed here. Without a well-defined challenge, even a sophisticated model produces nothing anyone can use.
Why it matters. A poorly defined problem produces misaligned expectations, wasted resources, and nothing of value at the end of it. A specific problem with measurable outcomes is the foundation everything else is built on.
A model is only as good as what trains it. Preparation is about more than cleanliness: the data has to align with the business problem and actually contain the patterns the question depends on.
Why it matters. Without relevant, high-quality data the initiative cannot succeed, whatever is spent on it. This step establishes whether the right data is in place before the commitment is made.
Consistent results come from a process, not from effort. Data preparation, model training and scaling all have to be repeatable — and the process needs a feedback loop, so the model keeps improving instead of quietly decaying.
Why it matters. A structured process is what makes output reliable rather than occasional, and it leaves room to refine and scale without starting again.
Prototype before committing. A prototype tests the assumptions made when the problem was framed and the data assessed — at the point where being wrong is still cheap.
Why it matters. Prototyping is how risk comes out of a programme. It is also the earliest point at which anyone outside the team can react to something real.
Invest the time upfront. Poorly framed problems are a leading cause of failure, and they surface as unmet expectations long after the budget is committed.
Gather and validate the data that supports the problem. It has to be relevant, reliable and representative of the domain — established before the work starts, not discovered during it.
Design toward the end goal: scalable, measurable, and matched to the resources available. Build for something that will change, and cut waste during development.
Set success criteria at every stage, not only at the end. Measurable goals are what let you track progress and manage expectations while the work is still in flight.
Set realistic accuracy expectations: perfection is not achievable. Measure against real-world benchmarks, and prefer a feasible solution to a theoretically better one.
Re-assess feasibility at each stage and plan resources accordingly. The project succeeds or fails on having the right people, tools and budget at the right time.
Make sure the solution handles the volume and complexity of real operations. Evaluate input types, processing needs and running cost before scale becomes the problem.
Outputs have to reach the business and be integrated into how it works. Plan how results are presented, so the people who make decisions can act on them.
These systems need maintenance. Design continuous feedback in from the start, so the solution improves on new data instead of drifting away from it.
Develop incrementally: prototype first, then build toward the full solution. Iteration is what refines the result and keeps risk contained.
Getting value out of AI and machine learning takes planning, clear objectives, and a structured route from problem to production. The sequence is the point: define the problem, establish the data, build the process, prove it small, then scale.
None of this is exotic. It is the unglamorous work that decides whether an investment returns anything — and it is the work most often skipped.
Riverton & Associates researches, designs, develops and implements production software and AI systems, using the approach set out here.
Riverton & Associates, Inc. · Austin, Texas · riverton-assoc.com
The Approach page sets out how these principles run as a delivery sequence on real engagements.
One supplier, forty-three spellings, and a finance team certain they know their top ten. Why entity resolution is the least glamorous and most valuable thing in the pipeline.
Every spend analysis rests on a question that sounds trivial: who did we pay? In practice the answer is scattered across dozens of spellings of the same name, in several systems, in records typed by people who had no reason to care about consistency.
Until those records are resolved to one supplier, every number downstream is wrong. Category totals are wrong. Negotiating leverage is understated. Duplicate payments hide. The top-ten supplier list — the one report every finance team trusts — is confidently, quietly incorrect.
Entity resolution is the work of turning many records into one supplier. It has no demonstration value, it produces no chart anyone wants to look at, and it is the single highest-return step in a spend pipeline.
One supplier generates variation from a dozen independent sources, and they compound.
Inc, Inc., Incorporated,
LLC, L.L.C., Ltd, GmbH,
Pty Ltd — with the suffix, without it, and punctuated three ways.MKTG and Marketing.
INTL, MFG, SVCS. Whoever set up the vendor
master had a forty-character field and used it.DELL 800-999-3355 TX. Same supplier, not one shared field
with the AP record.Forty-three spellings is not an exaggeration. It is what you get when a dozen sources of variation multiply across a few hundred thousand transaction lines.
The failure is not that the numbers are slightly off. It is that they are wrong in a direction nobody checks. Fragmentation always understates a supplier: spend that belongs to one relationship is divided among its variants, and each variant ranks lower than the whole.
| As recorded | Spend | Rank |
|---|---|---|
| DELL MKTG L.P. | $1.6M | 7 |
| Dell Marketing LP | $1.1M | 12 |
| DELL PRODUCTS L.P. | $0.8M | 19 |
| Dell Mktg L.P. (duplicate vendor ID) | $0.4M | 31 |
| card and expense lines | $0.2M | — |
| Resolved | $4.1M | 2 |
Illustrative. Every row is a real supplier with real invoices.
Nothing in the unresolved view suggests an error. The report is internally consistent, reconciles to the ledger, and is wrong about the thing it exists to tell you. The consequences are specific:
The instinct is to reach for fuzzy matching. It is not sufficient, and it is not even the main event.
High similarity, different entities. Dell Marketing L.P. and
Dell Products L.P. are almost identical strings and are genuinely
different legal entities with different contracts and payment terms. Whether they
should be merged depends on the question being asked — a business decision, not a
distance threshold.
Zero similarity, same entity. IBM and
International Business Machines Corporation. Alphabet and Google. Meta and
Facebook. No string metric will connect them. Only reference data will.
Name similarity is therefore one signal among several, and the others carry more weight than people expect: tax identifiers, registered address, bank remittance details, the GL codes a supplier is normally posted to, the buying entity, what was bought, and the behavioural fingerprint of invoice cadence and amount distribution. A match supported by four weak signals is stronger than one supported by a single strong name score.
Scale. Naïve comparison is quadratic — a few hundred thousand vendor records is tens of billions of pairs. The work has to be blocked, and blocking is where most of the recall is won or lost: a pair that never enters a block can never be matched, no matter how good the scorer is.
There is no single correct answer to who is this supplier, because the right grain depends on the question.
| Question | Right grain |
|---|---|
| What leverage do we have in a negotiation? | Ultimate parent |
| Who do we owe, and on what terms? | Legal entity |
| Who actually delivers, and what is our risk? | Operating entity or site |
| Which contract governs this line? | Contracting entity |
A layer that flattens everything to the ultimate parent cannot answer the second and third questions. One that stops at the legal entity cannot answer the first. The resolution is to keep the hierarchy rather than choose a level: resolve each record to a canonical legal entity with a stable identifier, then attach parentage above it, so reports roll up to whichever level the question needs. Choosing a single grain early is the most common irreversible mistake in this work.
False merges and false splits are not equally bad.
A false split leaves two records unmerged. The number is understated, the fragment is visible, and someone eventually notices a familiar name in the wrong place. It is a known unknown, and it is recoverable.
A false merge combines two genuinely different suppliers. The number is now confidently wrong, it reconciles, nothing looks unusual, and it has corrupted a figure someone is about to act on. It is very hard to detect after the fact, and it destroys trust in the whole layer when it is found.
This asymmetry should shape every threshold. Bias toward precision. Route the uncertain band to human review rather than guessing. And never let an automated merge be unexplainable: when a business owner challenges a number — and they will, usually the number that matters most — the answer has to be a specific reason, not a shrug.
Resist reporting percentage matched. It rewards over-merging, which is exactly the failure mode with the worst consequences. Better measures:
Set these criteria before building, not after. A resolution layer with no agreed measure of success cannot be said to have succeeded.
Entity resolution produces nothing anyone wants to demonstrate. There is no visualisation, no model architecture worth discussing, and no moment where it looks impressive. It is normalisation rules, blocking keys, calibrated thresholds and a review queue.
It is also the step on which everything else depends. Classification applied to fragmented suppliers classifies fragments. Anomaly detection on fragmented suppliers finds nothing, because the baseline is split across the variants. Every downstream number inherits the quality of this one.
The finance team is not wrong to trust their top ten. They are wrong about which suppliers are in it — and they have no way to know that from the report itself. Resolving the vendor name is what turns spend data into an answer somebody can defend.
Riverton & Associates implements the CXO Nexus enterprise spend analytics platform, including the resolution layer described here.
Riverton & Associates, Inc. · Austin, Texas · riverton-assoc.com
The CXO Nexus platform page shows where entity resolution sits in the pipeline, and what the record looks like on either side of it.
An old idea with relevance. Before the current wave, we captured how a thirty-year veteran made a judgement call and ran it at scale. That pattern is still the highest-return AI most operations can deploy.
Around twenty years ago we built an expert system with IBM for AK Steel. It encoded the judgement of a thirty-year veteran operator — decisions that had never been written down — and applied them on every shift rather than on the shifts he happened to work. It saved roughly $10 million a year.
It is still running. It is still saving money.
That fact deserves more attention than it usually gets. In an industry where a production model is doing well to survive eighteen months before drift forces a retrain, a system built two decades ago is still making the same decisions correctly. This paper argues that the pattern behind it — elicit what a skilled practitioner knows, encode it, run it everywhere — remains the highest-return application of AI available to most operations, and explains where it wins, where it loses, and why the historic barrier to building one has largely fallen away.
An expert system captures how a person makes a decision and executes that decision on every case, consistently, without fatigue.
It is worth being precise about the distinction from machine learning, because the two get confused. A learned model infers rules from labelled examples. An expert system takes the rules from a person who already has them. There is no training set, no labelling exercise, and no requirement that the past contains enough examples of the situation you care about.
That difference matters most where it is least convenient: rare events, new processes, and decisions that were never logged in a form a model could learn from.
A thirty-year veteran operator on a production line accumulates a body of knowledge that is almost entirely undocumented. He knows which combinations of conditions precede a problem. He knows when the specification is a hard limit and when it is a guideline. He knows which readings to trust and which sensor drifts. Most of this he cannot readily articulate — asked why he made a call, the honest answer is often because that is what it needed.
The value of encoding that was not that the system out-thought him. It was that it held the line when he was not there.
An operation does not run at the level of its best operator. It runs at the average of everyone on every shift, including the new hire on nights and the stand-in during holidays. The gap between the best operator's judgement and the median decision made across a year is enormous, and it shows up as scrap, rework, downtime and lost throughput. Closing that gap — not exceeding the expert, just reaching him consistently — was worth about $10 million a year.
The longevity is the part that should change how people think about this.
It is worth being honest about this rather than nostalgic. Expert systems fell out of use for real reasons, not only fashionable ones.
Then machine learning arrived, then deep learning, then language models, and each absorbed the attention and the budget. The pattern was not displaced because something better was found for the same problem. It was displaced because the industry's attention moved — and a technique that had stopped being interesting stopped being considered, while the systems already built carried on quietly working.
Fairness requires stating the other half. Expert systems are the wrong choice for perception — images, audio, free text — and for high-dimensional pattern recognition where the signal genuinely is a statistical regularity nobody could articulate. They are wrong where the domain really does drift, such as fraud and pricing, and they are wrong when nobody actually knows the rules. If your best practitioner cannot outperform chance, there is nothing to elicit and you need a model that learns.
The failure to avoid is treating this as an ideological choice. It is a question about the problem: does the knowledge exist in a person, and is the domain stable? If yes, encode it. If no, learn it.
Here is what has actually changed, and why the pattern deserves reconsideration now rather than merely respect.
The historic cost of an expert system was elicitation — the weeks of interviews, the analyst translating them into rules, the slow round trips to validate. That cost has fallen sharply. A language model can conduct the structured interview, work through decision logs and shift notes to propose candidate rules, spot contradictions between what an expert says and what the records show, draft the first implementation, and generate the test cases that prove it.
That suggests a specific architecture, and it is the one we would build today:
The model handles what it is genuinely good at: language, ambiguity, and interpretation. The decision itself stays in a component that is inspectable, testable, and cannot hallucinate. This gets the elicitation speed of the current wave with the durability and auditability of the old pattern — and it is a far better answer for an operations decision than putting a language model in the decision path and hoping.
Three reasons, and none of them are technical.
Expert systems did not fail. They went out of fashion, which is a different thing, and the industry stopped distinguishing between the two.
The system we built for AK Steel has been running for two decades and is still saving money — a claim very few modern AI deployments will be able to make in 2045. It did not out-think anyone. It took what one experienced person knew, and made it available on every shift.
That is still the most valuable thing most operations can do with AI. The tools for building it have improved considerably. The people whose judgement is worth capturing are retiring. Both of those facts point the same way.
Riverton & Associates researches, designs, develops and implements production software and AI systems, including the expert system described here.
Riverton & Associates, Inc. · Austin, Texas · riverton-assoc.com
If there is someone in your operation everybody defers to, that is the place to start — and the window is not open indefinitely.
Technology bought outside IT is not usually defiance. It is an expense line nobody categorized, in a system nobody joined to the asset register.
Most organisations discuss rogue IT — shadow IT, unsanctioned spend, whatever the local term is — as a behavioural problem. Somebody went around the process. The response is a policy memo, a stricter approval workflow, and a mild sense of grievance.
That framing is wrong, and it is why the problem never goes away.
Rogue IT is overwhelmingly a measurement failure. A manager needed a tool, the sanctioned route would have taken six weeks, and a corporate card took four minutes. The purchase was neither hidden nor defiant. It posted to a departmental expense code, under an approval threshold, against a merchant descriptor that matches nothing in the vendor master — and was never joined to the asset register, so no system anywhere knows the tool exists.
You cannot govern what you cannot see. Attack rogue IT first as a data problem: find it in the spend, quantify it, and only then decide what to do about it. The enforcement conversation is unproductive until the number exists.
Start by discarding the moral frame. In our experience the overwhelming majority of unsanctioned technology purchases share one profile: a competent person, doing their job, choosing the fastest route to a tool they needed.
If the sanctioned path takes six weeks and a credit card takes four minutes, the organisation has expressed a preference, whatever the policy says. People followed the incentive that was actually in place.
This matters practically, not just rhetorically. If you treat it as defiance, your response is enforcement — blocking, memos, tighter thresholds — which pushes the same purchases further out of view. Personal cards. Expenses coded as "training". Purchases made by a team in another country. You have not reduced the spend; you have reduced the visibility of it, which is the opposite of the goal.
The invisibility is mechanical, and each mechanism is mundane.
The direct overspend is real but it is rarely the largest number.
Only the first three are money. The rest are risk, and the risk is unquantified precisely because the spend was never measured.
Policy and blocking. Addresses the symptom, worsens the visibility, and does nothing about the six-week provisioning time that caused it.
Network discovery. Genuinely useful, and structurally incomplete. Traffic-based discovery misses anything used on a personal device, on a home network, or through a browser session it cannot inspect — which is most modern SaaS.
Asking people. Surveys of department heads return what they remember and are willing to declare. They often do not know either: the person who bought the tool left last year, and the card is now on someone else's expense report.
Each of these starts from the tool. The money is a better starting point, because a purchase always leaves a financial trace even when it leaves no technical one.
The reliable path runs through accounts payable and the card programme, not the network.
Lead with consolidation, not enforcement. The first pass usually reveals several departments paying separately for the same product. Consolidating those is a saving nobody has to be told off to achieve, and it buys the credibility for everything that follows.
Fix the risk before the commercials. Getting a discovered tool into SSO and the offboarding process is more urgent than renegotiating its price. Access outliving employment is the exposure that turns into an incident.
Fix the provisioning time. This is the actual root cause. If the sanctioned route stays six weeks, rogue IT regenerates continuously no matter how well you measure it. A fast-track path for low-value, low-risk tools removes most of the pressure.
Set thresholds that reflect annual value. Approve on the twelve-month cost, not the monthly charge, so a $400/month subscription meets the same scrutiny as a $4,800 purchase.
Rogue IT is not a story about people breaking rules. It is a story about an expense line that no system categorised as technology, attached to a supplier no system resolved, for a tool no system recorded.
Every part of that is a measurement failure, and every part is fixable with data work rather than policy. Resolve the supplier, classify the line, join it to the asset register, and the invisible spend becomes an ordinary number on an ordinary report.
Once it is a number, it can be managed. Until then, every conversation about it is a conversation about anecdotes.
Riverton & Associates implements the CXO Nexus enterprise spend analytics platform, which surfaces exactly this: technology spend resolved, classified at line-item level, and flagged against the cost centre that bought it.
Riverton & Associates, Inc. · Austin, Texas · riverton-assoc.com
The CXO Nexus platform resolves the supplier, classifies the line, and flags technology bought outside IT against the cost centre that bought it.
A model that wins on the metric and loses on the outcome is a common and expensive result. Choosing what to optimize is a business decision, not a technical one.
The most expensive failure in applied machine learning is not a model that performs badly. It is a model that performs well — against a target nobody should have chosen.
This failure is dangerous precisely because it does not look like failure. The metric improves. The validation curve behaves. The team ships on schedule and reports a win. Six months later the business number has not moved, and nobody can explain why, because every technical indicator is still green.
Choosing the objective is the highest-leverage decision in the whole project, and it is the one most often made by default — inherited from a tutorial, a benchmark, or whatever the library optimises out of the box.
Consider the shape of it. A team is asked to reduce customer churn. They build a churn classifier, tune it hard, and reach 0.91 AUC — a genuinely good model. It goes into production. Retention runs campaigns against its output. A year later, churn is unchanged.
Nothing in the modelling was wrong. The model predicts churn accurately. The problem is that accurate churn prediction and reduced churn are different objectives, and the project optimised the first while being paid for the second.
Worse, this failure resists debugging. Nobody investigates a model that is hitting its numbers. The instinct is to tune further — more features, a bigger model, better calibration — which improves the metric again and moves the outcome not at all. Teams can spend a year on that treadmill.
Every metric is a proxy for an outcome. Proxies diverge, and they diverge in specific, recognisable ways.
Underneath most of these is one omission: nobody stated what the errors cost. False positives and false negatives are almost never equally expensive, and the ratio between them determines the correct operating point. That ratio is a business fact. It cannot be derived from the data, and it is not the modeller's to invent.
| Churn example | Cost | What it is |
|---|---|---|
| Retention offer | $200 | what a false positive costs |
| Value of a saved customer | $3,000 | what a false negative costs |
| Ratio | 1 : 15 | — |
Illustrative. Ten minutes with the business owner establishes this.
At fifteen to one, you should be willing to accept a great many false positives to avoid one false negative. Yet the default instinct — and the default threshold in most libraries — is to balance the two, and a team asked to "improve precision" will tune away from exactly the customers worth saving. Establishing this ratio changes the threshold, the metric, and sometimes the entire choice of model. It is routinely skipped.
A prediction that changes nothing has no value, however accurate it is. Before choosing a metric, three questions should have concrete answers:
If those three cannot be answered, the metric cannot be chosen — and the honest response is to stop and answer them, not to pick a default and proceed.
When a measure becomes a target, it ceases to be a good measure. Not a new observation, but two things have changed.
First, optimisers are much stronger than they were. A modern model asked to maximise a proxy will find the gap between that proxy and the intended outcome, and it will find it efficiently. Capability makes this worse, not better — a weak model optimising the wrong objective fails harmlessly; a strong one succeeds at the wrong thing.
Second, the loop is faster. Systems retrain on data their own decisions generated, so a proxy mismatch compounds instead of staying constant.
The practical consequence: the more capable the model, the more precisely the objective needs to be stated. Vague targets were survivable when nothing could pursue them effectively.
This is the argument in the title, and it is worth stating plainly.
The technical team owns how to hit the target: architecture, features, training, validation, deployment. That is their expertise and it should not be second-guessed.
The business owns what the target is: which error is more expensive, what capacity exists to act, what horizon is useful, what must not get worse. These are not technical questions. They have no technically correct answer. They are statements about how the organisation makes money and what it is willing to trade.
When the objective is delegated to the technical team — usually by omission rather than decision — it defaults to whatever is conventional. And convention in machine learning is shaped by academic benchmarks, designed for comparability across papers, not for your profit and loss. Accuracy, F1 and AUC are excellent for ranking published results. None of them knows what a false negative costs you.
Any two of these together are enough to justify stopping and revisiting the objective. That conversation is uncomfortable — it implies the last several months optimised the wrong thing — and it is far cheaper than another two quarters of tuning.
Most failed machine learning projects do not fail technically. They hit their targets. The target was wrong, the wrongness was invisible because the metric kept improving, and the mistake was made in the first week by someone choosing a default.
Choosing what to optimise is the point where the business decides what it actually wants, in enough detail that a machine can pursue it relentlessly. That is not a task to delegate to whoever happens to be writing the training loop. It belongs to the people who know what the errors cost — and ten minutes of their time, before the modelling starts, is worth more than any amount of tuning afterwards.
Riverton & Associates researches, designs, develops and implements production software and AI systems, starting with what the system is actually for.
Riverton & Associates, Inc. · Austin, Texas · riverton-assoc.com
If the answer is not a sentence about who does what differently, that is the conversation to have before any modelling starts.
Designing an engagement so the client's team can run the system without us — and why that produces better architecture than building it to keep.
Most consultancies build to keep. Nobody says so, and it is rarely a conscious decision, but the incentives point that way and the resulting systems show it: bespoke conventions, tribal knowledge, a deploy that requires a phone call.
We design engagements the other way — so that the client's own team can run the system without us. That sounds like a commercial concession, and it partly is. But the substantive argument is different: designing for handover produces better architecture than designing to keep. Not as a happy side effect. As a direct consequence of the constraint.
Every requirement handover imposes — write it down, automate it, choose the boring option, make it observable — is something a well-built system should have anyway. Handover makes them non-optional, and supplies a test that cannot be argued with: someone who was not there has to be able to run it.
Building to keep does not usually look like sabotage. It looks like a series of individually reasonable decisions taken by people with no reason to decide otherwise:
None of this requires bad intent. It is what happens in the absence of a forcing function. Nobody sets out to build a system only they can run — it is the default outcome when nobody ever has to prove otherwise. The cost lands on the client eventually, and it is larger than the invoice: a system they depend on and cannot change, staffed by people they must keep hiring back.
Now impose one constraint: a competent engineer who was not here must be able to run this. Watch what that eliminates.
Read that list again with the handover framing removed. Every item is simply good engineering practice. Handover does not ask for anything that was not independently correct — it removes the option of skipping it, and supplies a test that produces a clear pass or fail.
That is the whole argument. The constraint is not a tax on the architecture. It is the quality bar, made externally verifiable.
The common pattern is to build for six months and then run a two-week "knowledge transfer" at the end. It reliably fails, for reasons that are structural rather than about effort.
Handover is a property of the system, not a stage of the project.
Honesty requires stating the commercial side. It is slower at the start — writing decisions down, automating the deploy before it is strictly needed, pairing rather than just doing it, all cost time in the early weeks.
It also forgoes lock-in. A client who can run the system without us can also choose not to call us. That is the point, and it does reduce a certain kind of recurring revenue.
What it produces instead is better work. Clients who can run what they own come back for the next problem rather than for maintenance of the last one — and the next problem is more interesting and more valuable than the maintenance. It also removes an objection at the point of sale: a buyer who fears dependency buys less, buys later, and buys with a smaller scope.
We would rather be invited back than needed.
The systems that come out of this discipline share recognisable traits:
Not it works in production.
Done is: their team shipped a change to it, without us, and we watched.
That is the acceptance criterion, and it is worth writing into the engagement. Anything short of it means the system is still on loan, regardless of what the status report says.
Designing an engagement for handover is not generosity, and framing it that way undersells it. It is the most reliable way we know to find out whether what we built is actually any good.
A system that cannot be handed over is one that has not finished being understood — its complexity is still being absorbed by the people who wrote it rather than resolved in the design. Making someone else run it is how you find that out, and doing so at the end is how you find it out too late.
Build it so they can run it without you. The architecture improves because of the constraint, not despite it.
Riverton & Associates researches, designs, develops and implements production software and AI systems — and designs the engagement so your team can run them.
Riverton & Associates, Inc. · Austin, Texas · riverton-assoc.com
If the honest answer is no, that is worth knowing now rather than at the end of the engagement.
A summary is a demo. An agent earns its place when it can take an action inside a system of record and be held to the result.
Most systems currently described as agents retrieve information, interpret it, and produce text. A person then reads that text and does the actual work. That is a useful assistant. It is not an agent, and the distinction is not pedantic — it is the difference between a system that demonstrates well and a system that changes a number somebody is accountable for.
The bar we hold to is specific: an agent takes an action inside a system of record, and somebody can be held to the outcome. Everything difficult about building one lives on the far side of that line, and nothing on the near side prepares you for it.
This paper is about what changes when a system stops describing the world and starts altering it: authority, reversibility, partial failure, audit, and the competence to decline.
A summarization demo is impressive in almost exact inverse proportion to what is at stake.
Nothing happened. No state changed, so there is nothing to reconcile, nothing to roll back, nobody to notify, and no audit trail to produce. The failure mode is a paragraph somebody disagrees with. The demo is safe because it is inert.
There is a second, subtler reason these demos flatter: summarization has no ground truth. A wrong summary looks exactly like a right one. It is fluent, plausible, and nobody in the room has read the underlying two hundred pages closely enough to catch the omission. Compare that to an agent that issues a credit note for the wrong amount — that failure announces itself, to someone, in a way nobody can talk around.
This is why pilots stall. The demo cleared a bar that had been set low by the absence of consequences, and the production version has to clear a different bar entirely.
Everything the demo omitted arrives at once, and none of it is model work.
None of these are improved by a better model. They are systems engineering, and they are where the actual work is.
The second half of the claim is the harder one. Accountability is what separates an agent from a suggestion engine, and it has concrete prerequisites.
If those five are not in place, the system may still be useful — but nobody is being held to anything, and its value is an assertion rather than a measurement.
The most important capability an acting agent has is knowing when not to act.
An agent that acts on everything is more dangerous than one that acts on sixty per cent and escalates the rest, because the sixty per cent it handles well tells you nothing about the forty per cent it should never have touched. Coverage is a poor target; coverage at an acceptable error rate is the real one.
That requires a calibrated sense of its own confidence, an explicit threshold below which it declines, and an escalation treated as a correct outcome rather than a failure. Teams that measure "percentage handled autonomously" will tune that threshold in exactly the wrong direction.
The related discipline is the cost asymmetry: how much worse is acting wrongly than not acting at all? For a reversible, low-value action the answer may be "barely". For a payment, a customer communication, or anything touching a regulated record, the answer is "enormously" — and the threshold should reflect that rather than a default.
The architecture that survives production separates interpretation from execution.
The model interprets. It reads the messy input, works out what is being asked, and proposes an action. This is what it is genuinely good at.
A deterministic layer executes. Actions are explicit, typed, individually permissioned tools — not free-form access to an API surface. Preconditions are checked in code before anything runs. Limits are enforced outside the model, where they cannot be argued with. Every execution writes its own audit record.
The model proposes; the deterministic layer disposes. This is the same division that makes expert systems durable, arriving in a modern context: the component that decides is inspectable, testable, and unable to invent a capability it does not have.
Practically, that means an agent should not have credentials. It should have a small set of narrow tools that have credentials, each of which validates its own preconditions and refuses outside them.
A summary is a demo. It is genuinely useful, it is easy to build, and it commits to nothing — which is exactly why it clears its bar so comfortably, and why so many of them never become anything more.
An agent earns its place when it changes something real: a record altered, a payment made, a case closed, a customer told. Everything hard about that is downstream of the moment it acts — authority, idempotency, reversal, audit, and the judgement to stop. None of it is model work, and all of it is the work.
Build the second thing. Then find out whether it was right, from a number, against a control, with somebody's name on it.
Riverton & Associates researches, designs, develops and implements production software and AI systems that take actions and are measured on them.
Riverton & Associates, Inc. · Austin, Texas · riverton-assoc.com
The useful place to start is one action, one owner, and a reversal rate you are willing to publish.
Riverton & Associates, Inc. is an Austin, Texas software and data consultancy. We research, design, develop, and implement systems that solve defined business problems — with deep expertise in data engineering, advanced analytics, model design, and architecting AI systems for production.
Applied well, it unlocks value in data that was previously impractical to reach: new products, better customer service, reduced risk, and materially more efficient operations. It enables smarter decisions, made faster.
Lack of clarity about the problem, poor execution, or goals misaligned with the business. Our methodology is built specifically to remove those failure modes rather than to demonstrate technology.
Systems in production at organizations where the results were measured. Several of this work also resulted in granted patents, but the patents are not the point — the systems are still running.
We work directly with CIOs, CTOs, and CEOs in financial services, insurance, and manufacturing — leaders accountable for productivity, competitive position, and the returns on a technology budget.
From problem definition through implementation and continuous improvement, rather than a strategy deck handed to somebody else to build.
Integrated with your people and your strategic vision, building in-house capability as the work proceeds.
A history of delivering measurable AI and machine learning value across regulated, operationally demanding industries.