Workflow Execution Engine
The Execution Engine is the heart of the Workflows module. It handles the actual running of workflows, managing state, error handling, and detailed logging.
Execution Overview
┌───────────────────────────────────────────────────────────────────────┐
│ EXECUTION ENGINE │
├───────────────────────────────────────────────────────────────────────┤
│ │
│ Trigger │
│ │ │
│ ▼ │
│ ┌──────────────────────────────────────────────────────────────┐ │
│ │ 1. Create WorkflowExecution record (status: pending) │ │
│ │ 2. Initialize execution context │ │
│ │ 3. Mark as running │ │
│ └──────────────────────────────────────────────────────────────┘ │
│ │ │
│ ▼ │
│ ┌──────────────────────────────────────────────────────────────┐ │
│ │ 4. Find entry point nodes (is_entry_point = true) │ │
│ │ 5. For each entry node: │ │
│ │ └─► executeNode(node, triggerData) │ │
│ └──────────────────────────────────────────────────────────────┘ │
│ │ │
│ ▼ │
│ ┌──────────────────────────────────────────────────────────────┐ │
│ │ executeNode(node, input): │ │
│ │ ├── Create WorkflowExecutionLog (status: pending) │ │
│ │ ├── Get node instance from NodeRegistry │ │
│ │ ├── Execute: node->execute(input, config, context) │ │
│ │ ├── Update log (status: completed, output data) │ │
│ │ ├── Get outgoing connections │ │
│ │ └── For each connection: │ │
│ │ ├── Check condition_path match │ │
│ │ └── Recursive: executeNode(nextNode, output) │ │
│ └──────────────────────────────────────────────────────────────┘ │
│ │ │
│ ▼ │
│ ┌──────────────────────────────────────────────────────────────┐ │
│ │ 6. Mark execution as completed/failed │ │
│ │ 7. Store final result │ │
│ │ 8. Fire completion events │ │
│ └──────────────────────────────────────────────────────────────┘ │
│ │
└───────────────────────────────────────────────────────────────────────┘
Execution Lifecycle
1. Triggering Execution
Workflows can be triggered in four ways:
| Trigger Type | How It Starts |
|---|---|
| Manual | User clicks Run button or calls POST /api/v1/workflows/{id}/execute |
| Schedule | Cron scheduler matches current time |
| Webhook | HTTP POST to POST /api/webhooks/workflows/{token} |
| Event | System event fires (e.g., order created) |
2. Creating Execution Record
When execution starts, a WorkflowExecution record is created:
WorkflowExecution::create([
'workflow_id' => $workflow->id,
'status' => 'pending',
'trigger_type' => 'manual', // or 'schedule', 'webhook', 'event'
'trigger_data' => $triggerData,
'triggered_by' => $userId,
'context' => [],
'nodes_total' => $workflow->nodes()->count(),
]);
3. Execution Context
The execution context is a shared state available to all nodes:
$context = [
'workflow_id' => 'uuid',
'execution_id' => 'uuid',
'trigger_data' => [ /* initial data */ ],
'variables' => [
// Set by SetVariable nodes
'key' => 'value',
],
];
Nodes access context via the third parameter:
public function execute(array $input, array $config, array $context): array
{
$workflowId = $context['workflow_id'];
$myVar = $context['variables']['myVar'] ?? null;
// ...
}
4. Node Execution
Each node execution follows this pattern:
protected function executeNode(WorkflowNode $node, array $inputData): array
{
// 1. Get node instance from registry
$nodeInstance = $this->nodeRegistry->make($node->node_type);
// 2. Create execution log
$log = WorkflowExecutionLog::create([
'execution_id' => $this->currentExecution->id,
'node_id' => $node->id,
'node_type' => $node->node_type,
'status' => 'pending',
'input_data' => $inputData,
]);
// 3. Execute the node
$log->markAsRunning();
$outputData = $nodeInstance->execute(
$inputData,
$node->config ?? [],
$this->executionContext
);
$log->markAsCompleted($outputData);
// 4. Follow connections
foreach ($node->outgoingConnections as $connection) {
if ($this->shouldFollowConnection($connection, $outputData)) {
$this->executeNode($connection->targetNode, $outputData);
}
}
return $outputData;
}
5. Conditional Branching
When a node outputs a _branch field, only matching connections are followed:
// Condition node output
{
"_branch": "true", // or "false", "case1", etc.
...otherData
}
// Connection has condition_path
// Only follow if _branch matches condition_path
if ($connection->condition_path === $outputData['_branch']) {
$this->executeNode($connection->targetNode, $outputData);
}
6. Completion
After all nodes execute, the execution is marked complete:
$execution->markAsCompleted($results);
// status: 'completed'
// result: final output data
// completed_at: timestamp
Execution States
Workflow Execution States
┌─────────┐
│ pending │
└────┬────┘
│ start
▼
┌─────────┐ delay or pause
┌───────│ running │───────────────────────┐
│ └────┬────┘───────┐ │
│ error │ success │ cancel ▼
▼ ▼ ▼ ┌────────┐
┌────────┐ ┌───────────┐ ┌───────────┐│ paused │
│ failed │ │ completed │ │ cancelled │└───┬────┘
└────────┘ └───────────┘ └───────────┘ │ resume
▲ │
└──────────────────────────┘
| Status | Description |
|---|---|
pending |
Created but not started |
running |
Currently executing |
completed |
All nodes finished successfully |
failed |
An error occurred |
cancelled |
Manually stopped |
paused |
Suspended part-way through, with its position saved |
Suspending and Resuming
A run does not have to hold a worker while it waits. Execution is checkpointed
after every node — the ready queue, unresolved join counts, pending inputs and
the step counter are written to workflow_executions.scheduler_state. A run can
therefore stop at any node boundary and be continued later without repeating work
whose side effects have already happened.
inputs holds only what nodes still to run are waiting on. Once a node has
executed and handed its output to every edge leaving it, the payloads it was
waiting on are spent and are dropped from the snapshot. That keeps the checkpoint
proportional to the graph's live frontier rather than to everything the run has
been through — which matters because the whole snapshot is rewritten after every
node, and nodes that echo their input make each stored payload a cumulative one.
The snapshot is format v2. v1 also stored every completed node's output
inline, and since the whole payload is rewritten after each node, the bytes a
run wrote grew with the square of its length — worse in practice, because nodes
commonly return array_merge($input, ...), so each stored output tended to
carry the previous one inside it. Those outputs were already being written to
workflow_execution_logs one row at a time, so a finishing run now rebuilds
them from there instead. v1 snapshots are still readable, so runs suspended
across the upgrade resume normally; support for them can be dropped once none
can still be in flight.
A node put back on the queue to be retried is the one case where an input is not dropped: it has produced nothing and has still to run, so the snapshot keeps what it was given and the retry resumes on the same data the first attempt saw. Pruning applies only to nodes that ran to completion.
Because of that, a queued node with no input in the snapshot is a contradiction
the engine will not paper over. Resuming from such a checkpoint fails with
Workflow node … is queued to run but its input is missing from the scheduler state, rather than running the node on an empty payload and quietly producing a
result built from nothing. Start a new execution if you see it.
The snapshot lives only as long as it is useful. It is discarded when a run
completes or is cancelled — neither is ever resumed, and the snapshot still
carries the payloads owed to whatever had yet to run, so keeping it would leave
that data on the bulk of the table for ever. A failed run deliberately keeps its
snapshot: the queue retries the job up to three times, and each retry resumes
from the checkpoint rather than re-running nodes whose side effects already
happened. Per-node history remains available either way in
workflow_execution_logs.
Three things suspend a run, and they differ only in whether a due time is set:
| Cause | status |
resume_at |
Resumed by |
|---|---|---|---|
Long Delay node |
paused |
when the wait elapses | workflows:resume-due |
| Long retry delay | paused |
when the next attempt is due | workflows:resume-due |
| Operator pause | paused |
null |
Re-dispatching the execution |
| Crashed or timed-out worker | running/failed |
null |
The queue's own retry |
resume_at is what keeps these apart: the sweeper only wakes executions that
have one, so a run paused by a person is never restarted behind their back.
Durable Delays
A Delay node longer than workflows.delay.pause_after_seconds (default 30)
does not sleep. It hands the wait back to the engine, which checkpoints the run,
records resume_at, and releases the worker. Shorter waits are still served in
process — the sweeper runs once a minute, so suspending a five-second delay
would stretch it towards a minute.
The scheduled command workflows:resume-due wakes due runs each minute, so a
delay elapses no earlier than requested and up to about a minute later. Design
around the lower bound, not the upper one.
Both scheduler commands walk every tenant that has the module enabled, since workflow tables live in per-tenant databases.
php artisan workflows:resume-due --dry-run # show what would be woken
Configuration
| Key | Default | Description |
|---|---|---|
workflows.delay.pause_after_seconds |
30 |
Waits longer than this suspend instead of sleeping. Keep it below the queue connection's retry_after. |
workflows.resume.claim_window_seconds |
300 |
How long a claimed run is hidden from other sweeps before falling due again. |
workflows.retry.pause_after_seconds |
30 |
Retry delays longer than this suspend the run instead of waiting in process. |
workflows.checkpoint.max_state_bytes |
1048576 |
Largest encoded scheduler_state the engine will store. Env: WORKFLOWS_MAX_STATE_BYTES. |
Checkpoint Size Ceiling
The snapshot is rewritten in full after every node, so its size is paid once per
step rather than once per run. Nodes commonly return array_merge($input, ...),
so a large payload entering the graph is carried forward and re-encoded for the
rest of it. workflows.checkpoint.max_state_bytes (default 1 MiB) is the
backstop against that growing without bound.
A run whose snapshot exceeds the limit is not failed — the queue retries a job three times, so failing a run that is otherwise succeeding would re-execute its nodes and duplicate their side effects. The snapshot is not truncated either: a shortened one still restores, into a run missing inputs it will never be told about. Instead the engine stops checkpointing that run and logs a warning naming the limit and the actual size. The run continues to completion normally.
The cost is that such a run can no longer be suspended. It will refuse an
operator pause and serve its Delay nodes in process, holding a worker for
them — the behaviour that existed before delays became durable. That is why the
default is generous: setting it too low quietly reintroduces worker-blocking
waits on ordinary large-payload workflows. If you see that warning, treat it as
a signal to trim what the workflow carries between nodes rather than to raise
the limit reflexively.
Operator Pause
POST /api/v1/workflow-executions/{execution}/pause asks a run to stop. The
worker is not interrupted: it finishes the node it is on, checkpoints, and stops
there, so the endpoint returns before the run has actually come to rest. Because
no resume_at is set, the run stays paused until it is explicitly re-dispatched.
A run whose checkpointing has failed repeatedly will refuse to pause and keep going — suspending a run whose position was never stored would leave it resumable in status only, with nothing to resume from.
Node Execution Log States
| Status | Description |
|---|---|
pending |
Node queued for execution |
running |
Node currently executing |
completed |
Node finished successfully |
failed |
Node threw an error |
skipped |
Node skipped (condition not met) |
Error Handling
Node-Level Errors
When a node fails:
- Exception is caught
- Execution log is marked as failed
- Error message and stack trace are stored
- Retry logic is checked
Retry Configuration
Nodes can be configured to retry on failure:
| Setting | Description |
|---|---|
retry_count |
Number of retry attempts (default: 0) |
retry_delay_seconds |
Delay between retries (default: 60) |
timeout_seconds |
Max execution time (default: 0 = unlimited) |
A node is run at most retry_count + 1 times.
Durable Retry Delays
A wait between attempts longer than workflows.retry.pause_after_seconds
(default 30) does not hold a worker. The node is put back on the ready queue,
the run is checkpointed and suspended with a due time, and the sweeper wakes it
exactly as it would a long Delay node. Shorter waits are still taken in
process.
The difference from a delay is what resuming does. A delayed node succeeded, so the run continues past it; a retrying node produced no output, so resuming runs that same node again.
Attempts already spent are recorded on the execution's context under
retries.{node_id}, so they survive the suspension. Without that the count
would restart on every resume, retry_count would never be reached, and a node
that always fails would retry indefinitely — one suspension at a time. The tally
is cleared once the node succeeds, and deliberately kept when it fails for good:
if the queue retries the whole job, the node has already spent its budget and
should not be handed a fresh one.
Each attempt taken after a suspension gets its own execution log row, stamped
with its retry_attempt number. Attempts retried in process share a single row
whose retry_attempt is incremented, as before.
Workflow-Level Errors
When any node fails and retries are exhausted:
- Execution is marked as failed
- Error node ID is recorded
WorkflowFailedevent is fired- Error message is stored
Execution Logs
Structure
Each node execution creates a log entry:
workflow_execution_logs
├── id: uuid
├── execution_id: uuid (FK)
├── node_id: uuid (FK)
├── node_type: string
├── node_label: string
├── status: enum
├── input_data: json
├── output_data: json (nullable)
├── error_message: text (nullable)
├── error_trace: text (nullable)
├── retry_attempt: integer
├── started_at: timestamp
├── completed_at: timestamp (nullable)
├── duration_ms: integer (nullable)
└── timestamps
Querying Logs
// Get all logs for an execution
$logs = WorkflowExecutionLog::where('execution_id', $executionId)
->orderBy('started_at')
->get();
// Get failed nodes
$failedLogs = WorkflowExecutionLog::where('execution_id', $executionId)
->where('status', 'failed')
->get();
// Get execution timeline
$timeline = WorkflowExecutionLog::where('execution_id', $executionId)
->select('node_label', 'status', 'started_at', 'duration_ms')
->orderBy('started_at')
->get();
Monitoring Executions
Viewing Execution History
Navigate to Workflows > [workflow] > Executions to see:
- List of all executions with status
- Trigger type and time
- Duration and node count
- Quick filters (status, date range)
Execution Detail View
Click an execution to see:
- Summary: Status, duration, trigger info
- Timeline: Visual progression through nodes
- Logs: Per-node input/output data
- Errors: Error messages with stack traces
Real-time Monitoring
For long-running workflows:
- Execution status updates in real-time
- Node progress indicator shows current step
- Logs appear as each node completes
Scheduled Execution
How Scheduling Works
-
Scheduler Command runs every minute:
php artisan workflows:run-scheduled -
Finds due workflows:
Workflow::where('status', 'active') ->where('trigger_type', 'schedule') ->get() ->filter(fn($w) => $this->isDue($w)); -
Dispatches jobs:
foreach ($dueWorkflows as $workflow) { ExecuteWorkflowJob::dispatch($workflow, [], 'schedule'); }
Cron Integration
Add to Laravel scheduler in app/Console/Kernel.php:
protected function schedule(Schedule $schedule)
{
$schedule->command('workflows:run-scheduled')
->everyMinute()
->withoutOverlapping()
->runInBackground();
}
Or the module automatically registers this when booted.
Webhook Execution
Webhook URL
Each workflow with webhook trigger gets a unique URL:
POST https://your-domain.com/api/webhooks/workflows/{workflow-token}
Signature Verification
For security, webhooks can require signature verification:
// Webhook node configuration
{
"requireSignature": true,
"signatureHeader": "X-Webhook-Signature",
"signatureSecret": "your-secret-key"
}
// Verification
$expectedSignature = hash_hmac('sha256', $payload, $secret);
$providedSignature = $request->header('X-Webhook-Signature');
if (!hash_equals($expectedSignature, $providedSignature)) {
abort(401, 'Invalid signature');
}
Webhook Response
Webhooks return immediately with execution ID:
{
"success": true,
"message": "Workflow execution started",
"data": {
"execution_id": "uuid",
"status": "pending"
}
}
Queue-Based Execution
Async Execution
For long-running workflows, use queue jobs:
// Dispatch to queue
ExecuteWorkflowJob::dispatch($workflow, $triggerData, 'manual', $userId);
// Job configuration
class ExecuteWorkflowJob implements ShouldQueue
{
public int $tries = 3;
public int $backoff = 60;
public int $timeout = 3600; // 1 hour max
public function handle(ExecutionEngine $engine)
{
$engine->execute(
$this->workflow,
$this->triggerData,
$this->triggerType,
$this->triggeredBy
);
}
}
Queue Configuration
In your .env:
QUEUE_CONNECTION=redis
Run the queue worker:
php artisan queue:work --queue=workflows
Events
Available Events
| Event | When Fired | Payload |
|---|---|---|
WorkflowStarted |
Execution begins | $execution |
WorkflowCompleted |
Execution succeeds | $execution |
WorkflowFailed |
Execution fails | $execution, $error |
Listening to Events
// In EventServiceProvider
protected $listen = [
\Modules\Workflows\App\Events\WorkflowCompleted::class => [
\App\Listeners\NotifyOnWorkflowComplete::class,
],
];
// Listener
class NotifyOnWorkflowComplete
{
public function handle(WorkflowCompleted $event)
{
$execution = $event->execution;
// Send notification, update stats, etc.
}
}
Performance Considerations
Large Workflows
For workflows with many nodes:
- Use queue-based execution
- Consider breaking into smaller workflows
- Monitor memory usage
High-Volume Triggers
For frequently-triggered workflows:
- Use dedicated queue workers
- Consider rate limiting
- Monitor queue depth
Long-Running Nodes
For nodes that take time (HTTP requests, delays):
- Set appropriate timeouts
- Use async patterns where possible
- Monitor execution duration
Debugging
Enable Debug Logging
// In config/workflows.php
'debug' => env('WORKFLOWS_DEBUG', false),
With debug enabled:
- All node inputs/outputs are logged
- Execution timing is recorded
- Stack traces are preserved
Common Issues
| Issue | Cause | Solution |
|---|---|---|
| Workflow not executing | Not activated | Set status to active |
| Node timeout | External service slow | Increase timeout_seconds |
| Missing data | Expression path wrong | Check {{expression}} paths |
| Infinite loop | Circular connections | Review workflow design |
Next Steps
- API Reference - Execution API endpoints
- Extending Workflows - Custom node execution
- Examples - Real workflow examples