Building Go's Goroutine Scheduler From Scratch in C
A stackful-coroutine scheduler + epoll netpoller in ~130 lines of C โ to understand what a goroutine actually is
The idea: I write Go’s
go func()every day but realised I couldn’t explain what a goroutine is under the hood, or why it’s cheap when an OS thread isn’t. So I rebuilt the core of Go’s runtime โ stackful coroutines scheduled on an epoll event loop โ from scratch in C. ~130 lines. It’s the clearest way I know to actually understand a tool: build the smallest working version of it.
A note on this write-up: the engineering and the learnings are mine; I used AI to turn my notes into readable prose. I’m a software engineer, not a writer, and I’d rather be upfront about that.
The question
A goroutine feels like magic: you launch a thousand of them, they block on I/O, and it just works โ cheaply. An OS thread can’t do that (a thread is ~1โ2 MB of stack and a syscall to schedule). So what is a goroutine? The honest answer is a stackful coroutine โ a function with its own small stack that can be paused and resumed โ plus a scheduler and a netpoller that parks it on I/O and wakes it when the I/O is ready. I wanted to build exactly that.
What I built
Two OS primitives do all the work:
ucontext(getcontext/makecontext/swapcontext) gives you a coroutine: a saved CPU context + a heap-allocated stack.swapcontext(a, b)saves where you are intoaand jumps intob. That’s a context switch, in user space, with no kernel involved.epollis the netpoller: register file descriptors, thenepoll_waitblocks until one is ready and tells you which.
The scheduler ties them together:
flowchart LR S[schedule loop] -->|swapcontext| C[coroutine runs] C -->|yields on I/O| S S -->|queue empty| E[epoll_wait blocks] E -->|fd ready| S S -->|re-queue that coro| C style S fill:#4a9eff,stroke:#4a9eff,color:#fff
schedule() drains a ready-queue, swapcontext-ing into each coroutine in turn. When a coroutine needs to do slow I/O, it registers its fd with epoll โ stashing its own Coro* pointer on the epoll registration โ and swaps back to the scheduler. When the ready-queue empties, the scheduler blocks in epoll_wait. An fd fires, epoll hands back the Coro* that was waiting on it, the scheduler re-queues it, and it resumes right after the swap, as if the blocking call had returned normally.
That handoff โ scheduler โ netpoller, with the waiting coroutine stashed on the fd โ is exactly how Go’s runtime parks and unparks goroutines on I/O. Building it by hand is what made goroutines stop being magic.
The payoff, proven
To prove it actually gives concurrency, I simulated slow I/O with timerfd (a timer that becomes a readable fd). Three coroutines each “do I/O” for 3s, 2s, and 1s โ and because each yields to the scheduler while waiting, they overlap on a single thread. Total wall time โ 3 seconds, not 6. One thread, no locks, genuine concurrency. That number is the whole point: it’s the event loop’s superpower, measured.
The first-principles insight: why C can’t have goroutines for free
Here’s the part that surprised me. You cannot bolt transparent async onto C the way Go has it โ and the reason is precise. Call it the libpq wall: when you call a blocking library function like PQexec, it blocks inside the library, on a socket you never get to see. You can’t register that socket with your epoll loop, so you can’t yield your coroutine โ the whole thread just stops.
Go and Node don’t win because their language is cleverer; they win because they own the network primitive. Go’s runtime is the netpoller, and its entire library ecosystem is written against it. That’s a whole-ecosystem property, not a syntax feature. Which leaves a real tradeoff triangle for a C server โ free-to-write blocking handlers ยท millions of connections ยท no runtime โ pick two.
This project was one point in a first-principles C series (a from-scratch HTTP server), alongside experiments in the other two concurrency models โ stackless coroutines (the Duff’s-device trick) and a thread-pool with an epoll wakeup channel โ so I could feel the tradeoffs between all three by building each.
Why this matters (beyond the fun)
It’s a learning artifact, not production code โ single-threaded, and CPU-bound work still needs real threads. But it changed how I talk about concurrency in interviews and design discussions: I no longer describe goroutines, I can explain the exact mechanism, because I’ve built the small version. Understanding beats memorising โ and the way I get understanding is to build the thing from first principles.