If you run any system that touches hiring, scheduling, discipline, or termination decisions, the ground just shifted. On September 30, 2026, Gov. Newsom signed SB 947 — the "No Robo Bosses Act" — which bars California employers from relying solely on automated decision‑making to fire or discipline workers. It also requires human review, disclosure of the data used, and a named human point of contact for affected employees. A CNBC report on the signing called it the first statewide human‑in‑the‑loop mandate for employment decisions in the U.S.
The mistake most teams make in the first week after a law like this lands: they treat it as a legal problem and hand it off to counsel. Legal can tell you what needs to be true. They can't build the checkpoint, wire the audit trail, or design the contested‑decision SLA. That part is operational, and it lands on you.
This playbook covers the runnable parts — the workflows, the artifacts, the owners, and the actual decisions you have to make so that "a human reviewed this" is genuinely true and provable, not a checkbox someone clicked at 4:59pm on a Friday.
The trap: "human‑in‑the‑loop" quietly becomes "human rubber‑stamp"
Most organizations already have some automation touching personnel actions. Scheduling tools that flag attendance. Productivity dashboards that surface "underperformers." Automated PIP triggers. Ticketing systems that route disciplinary flags to a manager's queue.
The failure pattern isn't that these systems make the final call. It's subtler than that. A manager gets a queue of 40 flagged items on Monday, each pre‑scored and pre‑justified by the system, and clicks "approve" on 38 of them in under ten minutes. Legally, a human was in the loop. Operationally, the human added nothing. The system decided; the human laundered the decision.
AI human‑in‑the‑loop compliance isn't satisfied by the presence of a human. It's satisfied by the presence of a human who had the information, the time, the authority, and the real option to disagree — and where you can prove all four later. That's the bar SB 947 actually sets, and it's the bar most workflows fail quietly.
The deeper problem the law exposes: a lot of teams never designed their automation to be reviewable. The model spits out a recommendation, the recommendation gets acted on, and nobody can reconstruct what inputs produced it. When an affected employee requests the data used — which they can now do — you discover the inputs weren't logged, the model version wasn't recorded, and the "human reviewer" field just says "system."
Map your decision surface before you touch anything
You can't add review checkpoints until you know where automated judgment actually enters personnel decisions. Most teams underestimate this surface by half, because they only count the obvious tools.
Stop losing track of your priorities.
Workyly helps you organize, assign, and track every task efficiently.
- Centralized task management
- Real-time collaboration
- Intelligent workflow automation
No credit card required
Start by inventorying every system that produces an output a manager could act on against a worker. Not just the ones labeled "AI." The attendance scoring. The sales leaderboard that auto‑ranks reps. The support QA tool that scores call transcripts. The gig‑style dispatch algorithm. If its output can reasonably influence someone getting disciplined, demoted, denied shifts, or fired, it's in scope.
| System | What it outputs | Does it influence a personnel action? | Current review before action |
|---|---|---|---|
| Attendance tracker | Absence/tardy score | Yes — feeds PIP triggers | None (auto‑escalates) |
| Support QA scorer | Call quality rating | Yes — used in reviews | Manager glances at score only |
| Shift dispatch engine | Shift allocation | Yes — low allocation = lost income | None |
| Productivity dashboard | "Flagged" employees | Indirectly — surfaces in 1:1s | Ad hoc |
The column that matters most is the last one. Where it says "none" or "glances at," you have a gap that SB 947 now treats as a liability. And the uncomfortable part: the "current review" column is only honest when someone other than the tool owner fills it out.
A practical tell — if you can't name the specific human who reviews a given system's output before it affects a worker, there is no human in that loop. There's a queue and a hope.
Build the review checkpoint so it can't degrade into a rubber‑stamp
A compliant checkpoint has to resist the Monday‑morning 38‑clicks problem. That means designing friction on purpose — not to slow everything down, but to make genuine review the path of least resistance.
-
The system produces a recommendation, never an action. No automated decision executes a personnel action directly. It creates a case, nothing more. This one rule eliminates the "the system fired someone" scenario entirely.
-
The case arrives with its reasoning and its inputs attached. The reviewer sees not just "flag: low performance" but the specific data points, the time window, the model version, and any obvious confounders the data doesn't capture. If a rep's numbers dropped the same week they were covering two territories, that context needs to be visible or at least discoverable.
-
The reviewer must record a rationale, not just an approval. A free‑text field with a minimum substance requirement. "Reviewed" doesn't count. "Confirmed three missed shifts, spoke with employee, no documented cause" does. This single requirement kills the ten‑minutes‑for‑40‑items pattern because you literally can't type 40 genuine rationales that fast.
-
Disagreement is a first‑class outcome. The interface should make "override — this recommendation is wrong" as easy as approving. If overriding requires a support ticket and approving is one click, you've built a rubber‑stamp whether you meant to or not.
-
Everything is logged immutably. Reviewer identity, timestamp, inputs shown, rationale, outcome. This is your proof later, and it's what you hand an employee who requests their data.
The non‑obvious upside here: the rationale requirement does double duty. It satisfies the human‑review mandate and it surfaces model quality problems you'd never catch from an approval‑rate metric alone. When reviewers keep overriding the same recommendation type, you've found a broken signal.
Require a minimum substance or a short template in the rationale field so approvals can't be one‑word placeholders.
This diagram gives a quick view of how the review flow should work.
The non‑obvious upside here: the rationale requirement does double duty. It satisfies the human‑review mandate and it surfaces model quality problems you'd never catch from an approval‑rate metric alone. When reviewers keep overriding the same recommendation type, you've found a broken signal.
The notice, data‑access, and point‑of‑contact machinery
The parts of SB 947 that teams tend to forget until it's too late are the employee‑facing ones. The law requires disclosure of the data used and a human point of contact for affected workers. According to the official release from Senator McNerney's office, these provisions put the employee's right to understand and contest decisions at the center of the law.
This is a data‑plumbing and SLA problem, not a policy‑statement problem.
For disclosure: when an automated system contributed to a decision, you need to produce what it used. That means logging inputs at decision time, not reconstructing them later from a database that's since been updated. A common mistake — teams log that a decision happened but query live tables to explain it, and the live data has changed. The disclosure you produce three weeks later doesn't match what the system actually saw.
For the point of contact: a real person, named, reachable, with authority to pull the case and the audit trail. Not a shared inbox that routes to nobody. Build a contested‑decision workflow with an actual SLA — acknowledge within X business days, provide the data within Y, route genuine disputes to a reviewer who wasn't the original one.
-
Every automated contribution to a personnel decision is disclosable in plain language
-
Inputs are snapshotted at decision time, not reconstructed from live data
-
A named human contact exists per decision type, with coverage for absences
-
A contested‑decision intake exists with a defined acknowledgment SLA
-
Contested cases route to a different reviewer than the original
-
Model version and data sources are stored with each case
-
Employees can request their data without going through the person who made the call
That last point matters more than it looks. If the only path to contest a decision runs back through the manager who made it, most people won't bother. The independence of the contest path is what makes the right real.
A real scenario: a mid‑size staffing operation
A regional staffing firm running roughly 600 field workers had an automated reliability score that fed directly into who got offered shifts. Low score, fewer shifts — which for a field worker is effectively a soft discipline, less income without anyone ever saying "you're being disciplined."
Before they addressed SB 947, nobody reviewed the scores. The algorithm pulled from clock‑in data, client ratings, and cancellation history, and the dispatch tool just allocated accordingly. When a worker complained they were getting no shifts, the ops team couldn't explain it beyond "the system."
They rebuilt it over about six weeks. The score became a recommendation surfaced to a dispatch supervisor, who reviewed any worker dropping below a threshold before allocation changed. They added rationale capture and a contact line. In the first month, supervisors overrode roughly 15% of the low‑score flags — mostly cases where a single bad client rating tanked an otherwise reliable worker, or where cancellations traced back to client‑side changes the algorithm had blamed on the worker.
The compliance piece was obvious. The quieter win: the override pattern showed that client ratings were far noisier than anyone had assumed. They reweighted the score accordingly. Shift‑allocation disputes dropped in the following quarter — not because of the law, but because human review exposed a broken input nobody had been watching.
When to pause a feature vs. when to wrap it in review
Not every automated system needs the same response. Some you harden with review; some you should just turn off until you can do it properly.
Wrap it in human review when: the automation produces genuinely useful signal, the inputs are logged or loggable, and a human can meaningfully evaluate the recommendation in a reasonable amount of time. Most scoring and flagging tools fit here.
Pause it when: the system currently executes personnel actions directly with no case step, the inputs can't be reconstructed for disclosure, or you can't name a reviewer who has the context to actually disagree. Running an unreviewable, unexplainable system during an enforcement window is worse than switching it off and doing the task manually for a few weeks.
Skip this entirely: adding a human "approver" purely as legal cover without changing how the work actually flows. That's the rubber‑stamp. And it's arguably more exposed than honest automation, because now you've documented a human taking responsibility for a decision they didn't actually make.
Why this becomes a governance problem, not a one‑time fix
The patch‑it‑once instinct is the real trap. You'll add checkpoints to the three systems you know about, pass your internal review, and six months later someone on another team ships a new "productivity insights" feature that quietly influences performance reviews — with no review layer, because nobody told them the rules applied to them.
Human‑in‑the‑loop compliance only holds if there's a standing process that catches new automation before it touches personnel decisions. An intake gate — any system that could influence a personnel action gets reviewed against the checkpoint requirements before it ships. Periodic re‑inventory too, because systems drift in scope. A dashboard that was "just informational" becomes decision‑shaping the moment managers start acting on it consistently.
This is the same discipline that governs any high‑stakes automation, and it's worth grounding your approach in a broader framework rather than a one‑law scramble. If you're building this muscle from scratch, the principles in our automation and AI governance framework for non‑engineering owners map cleanly onto what SB 947 now requires — ownership, review gates, logging, and a cadence for revisiting what your systems actually do.
The teams that handle this well won't be the ones with the best legal memo. They'll be the ones who already treated "a human can understand, review, and override this" as a design requirement — and for whom the law is a deadline, not a redesign. The ones who struggle will be the ones discovering, under audit, that their loop had a human in it the whole time who was never really deciding anything.
Start with the inventory. You can't review what you haven't mapped, and right now the biggest risk in most organizations isn't the system you're worried about — it's the two or three you forgot were making decisions at all.
Ready to boost your team's productivity?
Join 5,000+ teams using Workyly to streamline workflows, improve communication, and deliver projects faster.