| Problem | Solution | Result | |
|---|---|---|---|
| 01 | Monolithic server was hard to maintain and could not scale per feature | Migrated to MSA with Celery/Redis; switched worker pool prefork → gevent | Concurrent workflow slots per worker: 4 → 100 |
| 02 | Engine scanned the edge list on every step and ran nodes one at a time | Precomputed adjacency lists + dependency-based parallel scheduler | Next-node lookup O(E) → O(1); independent nodes run concurrently |
Moduly is an open-source AI automation platform where users build AI workflows by connecting nodes with drag and drop, then deploy and run them via API, webhook, or schedule.
A workflow is a directed graph of nodes (tasks) and edges (connections). The server's workflow engine starts at the start node and executes nodes by following the graph. Most nodes are I/O-bound tasks that spend most of their time waiting for a response, such as LLM or external API calls. Both challenges below stem from this characteristic.
The project started as a single FastAPI server, but two problems emerged as features grew.
I split the server into Gateway / Workflow Engine / Log System / Sandbox and chose Celery + Redis for communication between services.
At first I used the prefork (multi-process) pool. But each process held a single workflow and did nothing while waiting for an LLM response, and raising throughput meant adding more memory-expensive processes. The model did not fit I/O-bound work.
The alternatives, threads and gevent, differed in how well race conditions can be controlled.
| thread | gevent | |
|---|---|---|
| Scheduling | By the OS, preemptive | Cooperative, switches only at I/O points |
| Race conditions | Can occur at any point in the code | Possible points are limited to around I/O calls |
Both share memory, but gevent is never forcibly interrupted in the middle of a computation, which makes race conditions easier to reason about and control. On that basis I chose gevent.
During the migration, the existing asyncio-based engine caused event loop conflicts with the gevent pool, so I rewrote the engine and every node as synchronous code and unified the concurrency model on gevent alone.
concurrency): 4 → 100The workflow engine receives the graph only as a list of nodes and a list of edges.
When the engine is initialized, it scans the edge list only once to build a forward adjacency list (next-node lookup), a reverse adjacency list (predecessor lookup), and a per-branch index (path lookup for conditional branches). Every lookup during execution is then a single dictionary lookup (O(1)).
Some nodes can only run after their predecessors finish, but nodes with no dependency on each other can safely run at the same time. I built a scheduler around this.
To keep parallel execution safe, nodes only return their results and only the scheduler loop writes to the shared result store; submitted nodes are tracked to prevent duplicate execution.