Back to Blog

The SMS Communication Platform: Four Years of Production Hardening

· 7 min read
Share:
The SMS Communication Platform: Four Years of Production Hardening

Part of the Technical Breakdown series on my 4-year engagement with Prestige Men's Health.

Of everything I built for the clinic, the SMS platform is the one with the deepest history — over 340 commits across roughly three years, because it's the single busiest, most patient-facing, most real-time-sensitive system in the whole stack. Front desk staff live in it all day. Patients expect a text to show up instantly, not on the next polling cycle. That combination — high usage, real-time expectations, PHI-adjacent content — is exactly the kind of system that accumulates the most "found it in production, fixed it that afternoon" history. This post is that history.

The build: October–November 2023

The platform started from nothing in October 2023: a Twilio-backed two-way texting system with real-time delivery over Laravel Echo/WebSockets, so an inbound message appears in the front desk's conversation view the instant it arrives — no refresh, no polling. Standing it up also meant migrating roughly three years of the clinic's prior message history from their old messaging tool, attachments included, so staff didn't lose a single existing patient thread on cutover.

Problem: WebSocket connections silently going stale

Real-time delivery over a persistent WebSocket connection has a failure mode that doesn't show up in testing: a browser tab backgrounded for a while, a laptop coming back from sleep, a flaky office WiFi connection — the socket drops, but nothing in the UI tells you. Staff would report "I'm not getting new messages" with no error on screen, because from the browser's perspective, nothing had crashed — it just wasn't listening anymore.

// BEFORE — resources/js/sms-app.js (conceptual reconstruction)
Echo.channel(`user.${userId}`)
    .listen('NewSmsMessage', (e) => {
        appendMessageToConversation(e.message);
    });
// No reconnect logic — if the underlying socket drops, this listener
// is just gone. The page looks normal. Nothing fires again until a
// manual refresh.
// AFTER — resources/js/sms-app.js (conceptual reconstruction)
function subscribeToUserChannel(userId) {
    const channel = Echo.channel(`user.${userId}`)
        .listen('NewSmsMessage', (e) => appendMessageToConversation(e.message));

    return channel;
}

let channel = subscribeToUserChannel(currentUserId);

Echo.connector.pusher.connection.bind('state_change', (states) => {
    if (states.current === 'disconnected' || states.current === 'unavailable') {
        // Don't just hope Pusher/Echo's own retry handles it silently —
        // actively resubscribe once the underlying connection recovers,
        // and surface a visible indicator so staff aren't debugging a
        // "ghost" missing-messages report with no on-screen signal.
        showConnectionBanner('Reconnecting…');
    }

    if (states.current === 'connected') {
        hideConnectionBanner();
        channel = subscribeToUserChannel(currentUserId);
    }
});

// Browsers also throttle/suspend timers and sockets on backgrounded tabs —
// force a resubscribe check on visibility return, don't wait for the
// connection library to notice on its own.
document.addEventListener('visibilitychange', () => {
    if (document.visibilityState === 'visible') {
        channel = subscribeToUserChannel(currentUserId);
    }
});

This went through several iterations before it was solid — the real fix landed as "advanced websocket auto-reconnect" roughly two years after the initial build, and even that wasn't the end of it: browser-specific socket quirks (Edge, older Android WebViews) kept surfacing for months afterward. Real-time UI is one of those things where the happy path is a day of work and the reconnect/edge-case handling is a year of work.

Problem: Twilio media attachments not ready yet when the webhook fires

Twilio's inbound-MMS webhook can fire before the media asset is fully processed on their end — fetching the media URL immediately sometimes returned a transient error instead of the image.

// BEFORE
public function fetchMedia(string $mediaUrl): string
{
    $response = Http::get($mediaUrl);

    return $response->body(); // fails if Twilio hasn't finished processing yet
}
// AFTER
public function fetchMedia(string $mediaUrl, int $maxAttempts = 3): string
{
    $attempt = 0;

    do {
        $response = Http::get($mediaUrl);

        if ($response->successful()) {
            return $response->body();
        }

        $attempt++;
        usleep(300_000 * $attempt); // backoff: 300ms, 600ms, 900ms
    } while ($attempt < $maxAttempts);

    throw new MediaNotReadyException($mediaUrl);
}

Small fix, but it's the difference between "the patient's photo didn't come through" tickets and a system that just quietly retries for the roughly one second Twilio needs.

Problem: the same patient showing up as two separate conversations

Patients text in from a phone number before they're necessarily matched to an existing client record — a new inbound number needs to be reconciled against the clinic's patient list, and getting that matching wrong means the same person's messages scatter across duplicate contact records instead of one conversation thread.

// AFTER — app/Sms/Actions/ResolveContactForInboundNumber.php (conceptual)
final class ResolveContactForInboundNumber
{
    public function handle(string $phoneNumber): Contact
    {
        $normalized = PhoneNumber::normalize($phoneNumber);

        // Cascading match, most to least reliable signal:
        return Client::query()->where('phone', $normalized)->first()
            ?? Client::query()->where('email', $this->emailFromRecentForm($normalized))->first()
            ?? Client::query()
                ->where('first_name', $this->nameGuess($normalized)->first)
                ->where('last_name', $this->nameGuess($normalized)->last)
                ->where('date_of_birth', $this->dobGuess($normalized))
                ->first()
            ?? Contact::query()->firstOrCreate(['phone' => $normalized]);
    }
}

Even with that cascade in place at message-receive time, mismatches still slipped through at the margins — a typo'd digit, a patient texting from a new number before updating their file. The real fix wasn't just better matching logic, it was accepting that matching will never be 100% synchronous and adding a background reconciliation sweep: a scheduled job re-checks unmatched or ambiguous contacts every five minutes and merges them into the right client record retroactively, so a few minutes' lag beats a permanently split conversation history.

Problem: the conversation list getting slower as message volume grew

The front desk's conversation list — every active thread, most recent message preview, unread counts — got measurably slower as total message volume climbed into the hundreds of thousands. This went through multiple optimization passes over nearly two years as the dataset kept growing past whatever the last pass was tuned for.

// BEFORE (conceptual)
$conversations = Conversation::all();
foreach ($conversations as $conversation) {
    $conversation->last_message = $conversation->messages()->latest()->first(); // N+1
    $conversation->unread_count = $conversation->messages()->whereNull('read_at')->count(); // N+1 again
}
// AFTER (conceptual)
$conversations = Conversation::query()
    ->withMax('messages as last_message_at', 'created_at')
    ->withCount(['messages as unread_count' => fn ($q) => $q->whereNull('read_at')])
    ->with(['lastMessage:id,conversation_id,body,created_at'])
    ->orderByDesc('last_message_at')
    ->paginate(50);

Cursor-based pagination on top of this (rather than offset pagination) was the next necessary step once the list itself got long enough that "page 40 of conversations" became a real, requested thing — offset pagination gets slower the deeper you page, cursor pagination doesn't.

The AI layer: this platform became the agents' hands

By early 2026, the SMS platform wasn't just a messaging feature — it became the execution surface for autonomous AI tooling. There's a real app/Ai/Tools/ directory with purpose-built tools the AI agents call directly: SendPatientSmsTool, BulkSendGroupSmsTool, ScheduleSmsTool, SmsScopedSendLabDrawBookingLinkTool, SmsScopedSendPaymentLinkTool — each one a narrow, scoped action the agent can take, not open-ended database access. "Schedule a reminder text for this patient's next injection" or "send the lab-draw booking link to everyone flagged for a draw this week" became things an agent could actually execute, not just suggest.

This wasn't a single big-bang integration — it came out of roughly 65 autonomous GitHub Copilot coding-agent PRs merged in about six weeks (late February–early April 2026), each one a narrowly scoped change: copilot/optimize-sms-websocket-events, copilot/add-ai-sms-server-class, copilot/create-complex-ai-tools, copilot/fix-message-retrieval-logic, copilot/add-patient-lab-results-tool, copilot/fix-redis-cache-overload. That's the "autonomous AI engineering" phase from the main post, made concrete: dozens of small, reviewed, individually-mergeable PRs rather than one engineer hand-writing every line.

A small war story, because production always has one of these

My favorite bug from this entire project: a production 500 error traced down to a single missing space in a Blade template — @if with no space before it broke Blade's compiler in a way that only manifested after a specific conditional was added nearby. Hours of "but this worked yesterday" debugging for one whitespace character. Every long-running codebase has at least one of these; this was ours.


This is one of four deep-dives off the main technical breakdown post — the others cover the Kareo/Tebra EHR sync, the Rx order pipeline, and the patient portal.

Share:

Comments

No comments yet — be the first to share your thoughts.

Leave a comment

Your comment will be reviewed before it appears publicly.

Never published — only used if we need to reach you.

More from the blog