Technical Challenges and Solutions in the Moduly Project

edward·1일 전

💻 Click here to visit the github repository of MODULY

ProblemSolutionResult
01Monolithic server was hard to maintain and could not scale per featureMigrated to MSA with Celery/Redis; switched worker pool prefork → geventConcurrent workflow slots per worker: 4 → 100
02Engine scanned the edge list on every step and ran nodes one at a timePrecomputed adjacency lists + dependency-based parallel schedulerNext-node lookup O(E) → O(1); independent nodes run concurrently

About the Project

Moduly is an open-source AI automation platform where users build AI workflows by connecting nodes with drag and drop, then deploy and run them via API, webhook, or schedule.

  • Node types: LLM, knowledge base retrieval (RAG), HTTP request, Python code execution, conditional branch, loop, and more
  • Tech stack: Python 3.11 · FastAPI · Celery · Redis · PostgreSQL (pgvector) · Next.js · Docker · Kubernetes (EKS)

A workflow is a directed graph of nodes (tasks) and edges (connections). The server's workflow engine starts at the start node and executes nodes by following the graph. Most nodes are I/O-bound tasks that spend most of their time waiting for a response, such as LLM or external API calls. Both challenges below stem from this characteristic.


Challenge 01 — From Monolith to MSA, and Choosing a Worker Concurrency Model

Problem

The project started as a single FastAPI server, but two problems emerged as features grew.

  • Maintenance: As the server codebase grew, it became hard to see how far a change would reach.
  • Scaling: API handling, workflow execution, log ingestion, and code execution have different load characteristics, but because they lived in one process, we could not scale only the part that needed it.

Solution ① — MSA Migration and Asynchronous Communication

I split the server into Gateway / Workflow Engine / Log System / Sandbox and chose Celery + Redis for communication between services.

  • Loose coupling: The Gateway does not import engine code; it publishes work using only the task name.
  • Asynchronous processing: Workflow runs that take tens of seconds no longer block the API server from handling requests.
  • Independent scaling: Each service has its own image and Deployment, so only the service that needs it is scaled.

Solution ② — Celery Worker Pool: prefork → gevent

At first I used the prefork (multi-process) pool. But each process held a single workflow and did nothing while waiting for an LLM response, and raising throughput meant adding more memory-expensive processes. The model did not fit I/O-bound work.

The alternatives, threads and gevent, differed in how well race conditions can be controlled.

threadgevent
SchedulingBy the OS, preemptiveCooperative, switches only at I/O points
Race conditionsCan occur at any point in the codePossible points are limited to around I/O calls

Both share memory, but gevent is never forcibly interrupted in the middle of a computation, which makes race conditions easier to reason about and control. On that basis I chose gevent.

During the migration, the existing asyncio-based engine caused event loop conflicts with the gevent pool, so I rewrote the engine and every node as synchronous code and unified the concurrency model on gevent alone.

Result

  • Concurrent workflow slots per worker (Celery concurrency): 4 → 100
    • prefork: 4 processes, each holding one workflow. Raising it to 16 got the pod OOMKilled, so it was rolled back to 4.
    • gevent: 100 greenlets in a single process. The cost of one more concurrent workflow went from one process to one greenlet, so a workflow waiting on an LLM response no longer occupies a whole process.

Challenge 02 — Improving the Workflow Engine's Execution Algorithm

Problem

The workflow engine receives the graph only as a list of nodes and a list of edges.

  • Repeated computation: After executing a node, the engine scanned the entire edge list every time it looked for the next nodes.
  • Sequential execution: Even unrelated nodes were executed one after another.

Solution ① — Precomputed Adjacency Lists

When the engine is initialized, it scans the edge list only once to build a forward adjacency list (next-node lookup), a reverse adjacency list (predecessor lookup), and a per-branch index (path lookup for conditional branches). Every lookup during execution is then a single dictionary lookup (O(1)).

Solution ② — Dependency-Based Parallel Scheduling

Some nodes can only run after their predecessors finish, but nodes with no dependency on each other can safely run at the same time. I built a scheduler around this.

  1. Analyze each node's configuration for references to other nodes' outputs, and register only the nodes whose data is actually used as dependencies.
  2. When a node completes, look up the next-node candidates in the adjacency list.
  3. Among the candidates, submit as greenlets only those whose dependencies have all completed.
  4. Repeat until no nodes are running.

To keep parallel execution safe, nodes only return their results and only the scheduler loop writes to the shared result store; submitted nodes are tracked to prevent duplicate execution.

Result

  • Next-node lookup: full edge scan every time O(E) → O(1)
  • Execution time of a parallel section: sum of node times → time of the slowest node
profile
there ain't no shortcuts

0개의 댓글