Skip to content
Merged
Show file tree
Hide file tree
Changes from 6 commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion docs/ai/debugging.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,7 @@ These behaviors are especially hard to diagnose in a complex or long-running age

Durable workflows help by making it easier to **observe** the root cause of the failure, deterministically **reproduce** the failure, and **test or apply** fixes.
Because workflows checkpoint the outcome of each step of your workflow, you can review these checkpoints to see the cause of the failure and audit every step that led to it.
For example, using the [DBOS Console dashboard](../production/workflow-management.md), you might see that your agent failed because of a validation error caused by a malformed structured output:
For example, using the [DBOS Console dashboard](../conductor/workflow-management.md), you might see that your agent failed because of a validation error caused by a malformed structured output:

<img src={require('@site/static/img/why-dbos-agents/agent-fail.png').default} alt="Failing Agent" width="750" className="custom-img"/>

Expand Down
12 changes: 6 additions & 6 deletions docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -120,11 +120,11 @@ When operating DBOS durable workflows in production, we strongly recommend conne
Conductor is the control plane for your durable workflows, providing:

- [**High availability**](./production/workflow-recovery.md): In a distributed environment with many executors running durable workflows, Conductor automatically detects when the execution of a durable workflow is interrupted (for example, if its executor is restarted, interrupted, or crashes) and recovers the workflow to another healthy executor.
- [**Workflow and queue observability**](./production/workflow-management.md): Conductor provides dashboards of all active and past workflows and all queued tasks as well as real-time workflow visualization.
- [**Workflow and queue management**](./production/workflow-management.md): From the Conductor dashboard, you can pause any workflow execution, start any stopped or enqueued workflow, or restart any workflow from a specific step. This is useful for rapidly responding to incidents or debugging.
- [**Managed Retention Policies**](./production/retention.md): From the Conductor dashboard, manage how much workflow history each of your applications should retain and for how long to retain it.
- [**Autoscaling and version management**](./production/autoscaling.md): Conductor computes how many executors each version of your application needs from queue utilization, so autoscalers like KEDA can size a deployment per application version, drain old versions down to zero, and drive rollouts.
- [**Observability Integrations**](./production/metrics.md): Conductor exposes metrics about your applications' workflows, steps, and executors from a Prometheus-compatible endpoint, so you can monitor your DBOS applications in Datadog, Grafana, or any other tool that understands the OpenMetrics format.
- [**Workflow and queue observability**](./conductor/workflow-management.md): Conductor provides dashboards of all active and past workflows and all queued tasks as well as real-time workflow visualization.
- [**Workflow and queue management**](./conductor/workflow-management.md): From the Conductor dashboard, you can pause any workflow execution, start any stopped or enqueued workflow, or restart any workflow from a specific step. This is useful for rapidly responding to incidents or debugging.
- [**Managed Retention Policies**](./conductor/retention.md): From the Conductor dashboard, manage how much workflow history each of your applications should retain and for how long to retain it.
- [**Autoscaling and version management**](./conductor/autoscaling.md): Conductor computes how many executors each version of your application needs from queue utilization, so autoscalers like KEDA can size a deployment per application version, drain old versions down to zero, and drive rollouts.
- [**Observability Integrations**](./conductor/metrics.md): Conductor exposes metrics about your applications' workflows, steps, and executors from a Prometheus-compatible endpoint, so you can monitor your DBOS applications in Datadog, Grafana, or any other tool that understands the OpenMetrics format.

Architecturally, Conductor looks like this:

Expand All @@ -140,4 +140,4 @@ This architecture has two useful implications:
2. Conductor is **off your workflows orchestration path**. Conductor drives observability, recovery, and retention policies, and is never involved in workflow execution (unlike the external orchestrators of other workflow systems).
If your application's connection to Conductor is interrupted, it will continue to operate normally, and any failed workflows will automatically be recovered as soon as the connection is restored.

For more information on Conductor, see [its docs](./production/conductor.md).
For more information on Conductor, see [its docs](./conductor/overview.md).
4 changes: 2 additions & 2 deletions docs/production/alerting.md → docs/conductor/alerting.md
Original file line number Diff line number Diff line change
@@ -1,10 +1,10 @@
---
sidebar_position: 25
sidebar_position: 50
title: Alerting
toc_max_heading_level: 3
---

If you are using [Conductor](./conductor.md), you can configure automatic alerts when certain failure conditions are met.
If you are using [Conductor](./overview.md), you can configure automatic alerts when certain failure conditions are met.
You can configure alerts either in Conductor directly or on [Conductor-exported metrics](#metrics-based-alerts) using your existing observability stack.

:::info
Expand Down
Original file line number Diff line number Diff line change
@@ -1,10 +1,10 @@
---
sidebar_position: 32
sidebar_position: 80
title: Audit Logs
toc_max_heading_level: 3
---

If you are using [Conductor](./conductor.md), you can retrieve an **audit log** of the mutating operations performed against your organization: registering and deleting applications, managing workflows and schedules, creating and revoking API keys, changing roles and membership, and updating organization settings.
If you are using [Conductor](./overview.md), you can retrieve an **audit log** of the mutating operations performed against your organization: registering and deleting applications, managing workflows and schedules, creating and revoking API keys, changing roles and membership, and updating organization settings.
The audit log is append-only and records who did what, when, from where, and whether the operation succeeded.

:::info
Expand All @@ -13,7 +13,7 @@ Audit logs require a [DBOS Enterprise](https://www.dbos.dev/dbos-pricing) plan.

## The Audit Logs Endpoint

Conductor exposes an organization's audit log through the [Conductor API](./conductor-api.md) at:
Conductor exposes an organization's audit log through the [Conductor API](./reference/conductor-api.md) at:

```
GET https://cloud.dbos.dev/conductor/v2/orgs/{orgName}/audit-logs
Expand All @@ -34,7 +34,7 @@ curl -G https://cloud.dbos.dev/conductor/v2/orgs/my_org/audit-logs \
```

:::note
Audit logs are an organization-level concept, so a [self-hosted Conductor](./hosting-conductor.md) running with authentication disabled does not register this operation and responds `404`. See [Self-hosted differences](./conductor-api.md#self-hosted-differences).
Audit logs are an organization-level concept, so a [self-hosted Conductor](./self-hosting/hosting-conductor.md) running with authentication disabled does not register this operation and responds `404`. See [Self-hosted differences](./reference/conductor-api.md#self-hosted-differences).
:::

Entries are returned newest first (by emit time).
Expand Down
Original file line number Diff line number Diff line change
@@ -1,10 +1,10 @@
---
sidebar_position: 22
sidebar_position: 60
title: Autoscaling and Version Management
toc_max_heading_level: 3
---

[Conductor](./conductor.md) lets you attach autoscaling policies to your applications. An autoscaling policy computes how many executors your application needs, per application version, to drain one of your application's queues. A common example is configuring a [KEDA](https://keda.sh/) ScaledObject to size your application deployments based on queue utilization.
[Conductor](./overview.md) lets you attach autoscaling policies to your applications. An autoscaling policy computes how many executors your application needs, per application version, to drain one of your application's queues. A common example is configuring a [KEDA](https://keda.sh/) ScaledObject to size your application deployments based on queue utilization.

:::info
Autoscaling requires a [DBOS Teams](https://www.dbos.dev/dbos-pricing) plan.
Expand All @@ -15,7 +15,7 @@ To use policies:
1. **Attach an autoscaling policy** to an application, naming the queue whose backlog drives the executor count.
2. **Poll the desired executor count**, either one version at a time or for all active versions at once.

All endpoints on this page are part of the [Conductor API](./conductor-api.md); see that page for the base URL and authentication.
All endpoints on this page are part of the [Conductor API](./reference/conductor-api.md); see that page for the base URL and authentication.
The examples below use `$CONDUCTOR` for the base URL and `$CONDUCTOR_KEY` for an [API key](./permissions.md).

## How It Works
Expand Down
16 changes: 16 additions & 0 deletions docs/conductor/distributed-recovery.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
---
sidebar_position: 30
title: Distributed Recovery
---

If your application is connected to [DBOS Conductor](./overview.md), workflow recovery is automatic.
When Conductor detects that an executor is unhealthy, it automatically signals another executor to recover its workflows.

When an executor disconnects from Conductor, its status is changed to `DISCONNECTED` while Conductor waits for it to reconnect.
If it has not reconnected after a certain period of time, its status is changed to `DEAD` and Conductor signals another executor to recover its workflows.
After recovery is confirmed, Conductor deletes its record of the executor.

By default, the executor timeout is 60 seconds, so Conductor waits 60 seconds after an executor disconnects before recovering its workflows.
You can configure the executor timeout per application from the DBOS Console.

<img src={require('@site/static/img/conductor/grace-period.png').default} alt="Workflow Timeout Configuration" width="800" className="custom-img"/>
4 changes: 2 additions & 2 deletions docs/production/metrics.md → docs/conductor/metrics.md
Original file line number Diff line number Diff line change
@@ -1,10 +1,10 @@
---
sidebar_position: 24
sidebar_position: 40
title: Metrics
toc_max_heading_level: 3
---

If you are using [Conductor](./conductor.md), you can scrape metrics about your applications' workflows, steps, and executors from a [Prometheus](https://prometheus.io/)-compatible endpoint.
If you are using [Conductor](./overview.md), you can scrape metrics about your applications' workflows, steps, and executors from a [Prometheus](https://prometheus.io/)-compatible endpoint.
This lets you monitor your DBOS applications in Prometheus, Grafana, or any other tool that understands the [OpenMetrics](https://prometheus.io/docs/specs/om/open_metrics_spec/) format.

:::info
Expand Down
8 changes: 4 additions & 4 deletions docs/production/conductor.md → docs/conductor/overview.md
Original file line number Diff line number Diff line change
@@ -1,18 +1,18 @@
---
sidebar_position: 10
title: DBOS Conductor
sidebar_position: 1
title: DBOS Conductor Overview
---

When operating DBOS durable workflows in production, we strongly recommend connecting your application to Conductor.
Conductor is the control plane for your durable workflows, providing:

- [**High availability**](./workflow-recovery.md): In a distributed environment with many executors running durable workflows, Conductor automatically detects when a workflow is interrupted (for example, if its executor disconnects or crashes) and recovers the workflow to another healthy executor.
- [**High availability**](./distributed-recovery.md): In a distributed environment with many executors running durable workflows, Conductor automatically detects when a workflow is interrupted (for example, if its executor disconnects or crashes) and recovers the workflow to another healthy executor.
- [**Workflow and queue observability**](./workflow-management.md): Conductor provides dashboards of all active and past workflows and all queued tasks as well as real-time workflow visualization.
- [**Workflow and queue management**](./workflow-management.md): From the Conductor dashboard, you can pause any workflow execution, start any stopped or enqueued workflow, or restart any workflow from a specific step. This is useful for rapidly responding to incidents or debugging.
- [**Managed Retention Policies**](./retention.md): From the Conductor dashboard, manage how much workflow history each of your applications should retain and for how long to retain it.
- [**Autoscaling and version management**](./autoscaling.md): Conductor computes how many executors each version of your application needs from queue utilization, so autoscalers like KEDA can size a deployment per application version, drain old versions down to zero, and drive rollouts.
- [**Observability Integrations**](./metrics.md): Conductor exposes metrics about your applications' workflows, steps, and executors from a Prometheus-compatible endpoint, so you can monitor your DBOS applications in Datadog, Grafana, or any other tool that understands the OpenMetrics format.
- [**Programmatic access**](./conductor-api.md): Conductor's workflow, queue, and schedule management is available over an OpenAPI-described HTTP API and from the [`dbosctl` command-line client](./dbosctl.md), so you can script incident response and wire Conductor into your own tooling.
- [**Programmatic access**](./reference/conductor-api.md): Conductor's workflow, queue, and schedule management is available over an OpenAPI-described HTTP API and from the [`dbosctl` command-line client](./reference/dbosctl.md), so you can script incident response and wire Conductor into your own tooling.

Architecturally, Conductor is not part of your workflows orchestration path.
If your connection to Conductor is interrupted, your applications will continue operating normally.
Expand Down
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
---
sidebar_position: 31
sidebar_position: 70
title: Permissions and API Keys
---

Expand Down Expand Up @@ -76,7 +76,7 @@ Like a role, every API key carries a set of permissions.
They can also be scoped to specific applications.

API keys do not expire, but can be revoked at any time.
A key can be renamed after creation without changing its secret, from the console, with [`dbosctl api-key rename`](./dbosctl.md#dbosctl-api-key-rename), or through the [Conductor API](./conductor-api.md#roles-permissions-and-api-keys).
A key can be renamed after creation without changing its secret, from the console, with [`dbosctl api-key rename`](./reference/dbosctl.md#dbosctl-api-key-rename), or through the [Conductor API](./reference/conductor-api.md#roles-permissions-and-api-keys).

### Permissions and application scope

Expand All @@ -89,7 +89,7 @@ For example, an API key with only `application.read` scoped to a single applicat

### Using an API key

Supply the key to your DBOS application to connect it to Conductor, as described in [Connecting to Conductor](./conductor.md#connecting-to-conductor).
Supply the key to your DBOS application to connect it to Conductor, as described in [Connecting to Conductor](./overview.md#connecting-to-conductor).

You can also use an API key to authenticate HTTP calls to the Conductor API (for example the [metrics endpoint](./metrics.md)), passing the key as a bearer token:

Expand Down
4 changes: 4 additions & 0 deletions docs/conductor/reference/_category_.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
{
"label": "Reference",
"position": 100
}
Loading
Loading