A marketplace held real money in escrow and promised to release it automatically after 24 hours. The job had been deployed for 96 days. It had never run once.
The platform is a two-sided marketplace connecting funeral homes with verified workers — removals, embalming, directing, visitation staff — plus vehicle and equipment suppliers. Money moves through it. A funeral home's payment is captured and held in escrow when a booking is made, and released to the worker after the job is marked complete and a 24-hour dispute window closes.
That release was supposed to be automatic. The database functions existed. The status column flipped from held to released exactly as designed. Everything about the system said the payout had happened.
No money had moved. The functions updated the payment status without ever calling Stripe, and nothing was scheduled to invoke them in the first place. On a pre-launch platform with no live volume yet, nothing surfaced the gap — the dashboards were reporting a state that had no transfer behind it.
This surfaced during a paid full-lifecycle test — deliberately walking the entire path from post through payout rather than testing features in isolation:
Two problems came out of that walk. The first was an authorization gap: administrators had no SELECT policy on bookings, worker profiles, funeral home profiles, or jobs, so the admin manual payment-release override returned 404 before its logic ever ran. Fixed in a migration and confirmed with a real test-mode transfer.
The second was the auto-release itself. I rebuilt it as an hourly Vercel cron hitting a dedicated route that performs the actual Stripe transfer, guarded by a shared secret and keyed with a per-booking Stripe idempotency key so a retry can never double-pay a worker. Verified with a script that drove the whole path. Shipped.
Three days later it still had not fired.
The first theory was a missing CRON_SECRET in the production environment — the guard rejecting its own scheduler. Plausible, easy to believe, and wrong: the secret had been set the entire time, 96 days.
The second theory was the hosting plan. Hourly cron schedules are a paid-tier feature, so a project sitting on a lower tier would silently get a coarser schedule or none at all. Also plausible. Also wrong — the project was on Pro, and hourly had always been supported.
The actual cause was one line of middleware. The platform was in a pre-launch lock, with every route redirected to the holding page unless explicitly whitelisted. The whitelist covered the Stripe webhook path. It did not cover /api/cron/*. So the scheduler dutifully fired every hour, received a 307 redirect to the marketing page before the handler was reached, treated the redirect as a successful invocation, and moved on. Nothing in the code review would have caught it. Nothing in the logs looked like an error.
The fix was whitelisting the cron prefix. The part worth keeping was the verification: the route now returns its own 401 rather than a 307, a request with a wrong bearer token is still refused, and the admin dashboard still redirects — proving the fix opened exactly one door and left the pre-launch lock otherwise intact.
Both of my first two explanations were coherent, mechanical, and consistent with the symptom. Neither was checked before it was believed. The thing that actually resolved it was probing the deployed endpoint and reading what it returned — five minutes of looking at the system instead of reasoning about it.
A deployed job is not a running job. A status column is not a payment. Now the standard on anything that moves money or runs unattended is the same: watch it work against the real system, confirm the negative case is still refused, and treat a plausible mechanism as a hypothesis until an observation makes it a finding.
Hiring, contract work, or just a question about how this one was built — email is the fastest path.
FormationLabs
AI Assistant
Quick questions: