Last week, I tagged Francis 0.1.0, the first stable release of the distributed actor framework for Go I’ve been building over the last 18 months.
Francis is a framework and runtime for distributed actors (also known as durable objects), and a durable workflow engine built on top of them. An actor is a unit of state with single-threaded compute on top, addressed by a type and an ID: you write it as a plain Go struct, and Francis activates it on demand somewhere in the cluster, routes calls to it, persists its state, and runs its alarms. If the actor model is new to you, I wrote a longer introduction last year. (As for the name, I ended up keeping the “tentative” one from when I wrote about designing it back in January.)
While the “stable” label is new, Francis has been running in production since July inside Pocket ID, the self-hosted OpenID Connect provider that lets users sign in with passkeys. Pocket ID has an estimated 10,000+ installations, which, as far as I can tell, makes Francis one of the most widely deployed actor frameworks for Go already (and probably other languages too).
Where Francis comes from
Much of Francis builds on what I learned before I started writing it. I was a maintainer of Dapr, which includes a distributed actor runtime (Dapr Workflow is built on top of those actors, too), and I’ve used Orleans, the .NET framework that popularized the virtual actor model.
Orleans and Dapr sit at opposite ends of a spectrum. In Orleans, the runtime is a library that runs inside your process, so your app itself is the actor host. That makes for a great developer experience, but only if your app is written in .NET. Dapr puts the runtime in a sidecar next to your app, and actors also depend on control plane services (the placement service and, in recent versions, the scheduler) that you have to deploy and operate. Dapr is language-agnostic and well suited to large clusters, but it requires multiple control plane services and it’s overly complex for smaller apps that are not (yet) scaling horizontally.
Francis borrows from both. Like Orleans, it’s a library you import into your app, with an API that feels native to Go. Like Dapr, it can also run its control plane as a separate service, but only when your deployment calls for it: in its simplest form, all Francis needs is a relational database, either SQLite or PostgreSQL.
Scaling down as well as up
The core idea behind Francis is that an actor framework should scale down as well as it scales up.
Most actor frameworks assume you already have a distributed system: several replicas of services that could be heterogeneous and need coordination. However, the problems actors solve show up long before that. Many apps I’ve worked on have needed some background work on a schedule, some rate limiting, or a multi-step process that has to finish even if the server restarts halfway through. Common solutions keep timers and state in memory, which works until you add a second replica, then requiring complex state coordination.
With Francis, an app running as a single replica with a SQLite file gets durable alarms, jobs, workflows, and ready-to-use built-in actors such as a rate limiter. When the app grows to a few replicas, the same code keeps working unchanged. Francis makes sure each actor is active on exactly one host at a time, so a scheduled job still runs once, and each rate limit is still counted in one place. As the solution keeps growing into dozens of instances across heterogeneous services, you can move coordination into a dedicated runtime, and your actors don’t change.
Francis was designed to support two topologies:
- Local: Francis is embedded in your Go app, and there’s no other service to run. Actor state and alarms live in SQLite or PostgreSQL (which can be the same database your app already uses), and when there’s more than one replica, hosts coordinate peer-to-peer through the shared database. This is ideal for a single instance or a small cluster (up to about 4 replicas).
- Remote: a standalone runtime, which can itself run as a cluster of replicas, owns the database and coordinates placement, state, and alarms. Your apps become stateless hosts that connect to it and scale independently. This is the topology for larger clusters, auto-scaling, and multiple services that host and invoke each other’s actors.
Your actor code is identical in both. The only thing that changes is how you construct the host:
func newHost(runtimeAddresses []string) (host.Host, error) {
// No runtime configured: run Francis embedded in the app
if len(runtimeAddresses) == 0 {
return local.NewHost(
local.WithAddress("127.0.0.1:7571"),
local.WithSQLiteProvider(local.SQLiteProviderOptions{
ConnectionString: "data.db",
}),
local.WithRuntimePSKs(runtimePSK),
)
}
// Otherwise, connect to a standalone runtime (or a cluster of them)
return remote.NewHost(
remote.WithAddress("127.0.0.1:7571"),
remote.WithRuntimeAddresses(runtimeAddresses...),
remote.WithHostBootstrapPSK(hostBootstrapPSK),
remote.WithPinnedCA(clusterCA),
)
}
Pocket ID was the first user of Francis (outside of apps I manage for myself), and in July it moved its background jobs, cron jobs, and rate limiting to Francis. Pocket ID doesn’t support running more than one replica yet, so those installs run as a single container, with Francis embedded and storing its data in Pocket ID’s own database (SQLite by default, or PostgreSQL). Thanks to Francis, the upcoming Pocket ID v3 will (finally) support running multiple replicas, which has been a longstanding feature request.
A quick tour
This is only a quick overview: full details and more examples are in the docs.
Actors
An actor is a Go struct, plus a factory function that Francis calls whenever it needs to activate one. The interfaces the struct implements determine what it can do: Invoke handles method calls, while Alarm and Job receive (you guessed it) alarms and jobs, respectively.
This is a counter that keeps its value in durable state:
type Counter struct {
client actor.Client[counterState]
}
// counterState is the actor's durable state
type counterState struct {
Count int64
}
// NewCounter is the factory Francis calls to activate a Counter
func NewCounter(actorID string, service *actor.Service) actor.Actor {
return &Counter{
client: actor.NewActorClient[counterState]("counter", actorID, service),
}
}
func (c *Counter) Invoke(ctx context.Context, method string, data actor.Envelope) (any, error) {
// Returns the zero value if the actor has no state yet
state, err := c.client.GetState(ctx)
if err != nil {
return nil, err
}
switch method {
case "increment":
state.Count++
case "reset":
state.Count = 0
default:
return nil, fmt.Errorf("unknown method: %s", method)
}
err = c.client.SetState(ctx, state, nil)
if err != nil {
return nil, err
}
return state.Count, nil
}
To run it, create a host, register the actor type, and invoke actors through the host’s service. This uses the local topology with a SQLite database, which is the most barebones setup:
h, err := local.NewHost(
local.WithAddress("127.0.0.1:7571"),
local.WithSQLiteProvider(local.SQLiteProviderOptions{
ConnectionString: "data.db",
}),
// The cluster CA, used for mTLS between hosts, is derived from this key
local.WithRuntimePSKs([]byte("change-me-please")),
)
if err != nil {
log.Fatal(err)
}
err = h.RegisterActor("counter", NewCounter)
if err != nil {
log.Fatal(err)
}
// Run blocks until the context is canceled, so start it in the background
go func() {
err := h.Run(ctx)
if err != nil {
log.Fatal(err)
}
}()
<-h.Ready()
// Invoke the "increment" method on the counter with ID "alice"
resp, err := h.Service().Invoke(ctx, "counter", "alice", "increment", nil)
if err != nil {
log.Fatal(err)
}
var count int64
err = resp.Decode(&count)
Each actor ID is a separate actor with its own state, so alice and bob count independently, and because state is persisted in the database, the counts survive restarts. Calls to the same actor are processed one at a time, so there’s no need for a mutex around the read-modify-write.
Actors can also schedule alarms for themselves, which Francis persists and delivers even if the actor was hibernated or the process restarted. For example, a shopping cart can use this to garbage collect itself after inactivity (pushing the deadline forward every time something is added):
func (c *Cart) Invoke(ctx context.Context, method string, data actor.Envelope) (any, error) {
switch method {
case "add-item":
// ... update the cart's state ...
// Setting an alarm with the same name replaces it, which resets the deadline
return nil, c.client.SetAlarm(ctx, "expire", actor.AlarmProperties{
DueTime: time.Now().Add(24 * time.Hour),
})
}
return nil, nil
}
func (c *Cart) Alarm(ctx context.Context, name string, data actor.Envelope) error {
if name == "expire" {
// ... delete the cart ...
}
return nil
}
Jobs
Jobs are durable, fire-and-forget tasks dispatched to an actor. Jobs that fail are retried, and when they run out of attempts they’re dead-lettered, so you can inspect and replay them later.
An actor receives jobs by implementing Job:
func (u *User) Job(ctx context.Context, method string, data actor.Envelope) error {
switch method {
case "send-welcome-email":
var req welcomeEmail
err := data.Decode(&req)
if err != nil {
// Bad input won't fix itself: skip the retries and dead-letter the job right away
return errors.Join(actor.ErrJobPermanentFailure, err)
}
return u.sendWelcomeEmail(ctx, req)
case "send-weekly-summary":
return u.sendWeeklySummary(ctx)
}
return nil
}
Any part of your app can dispatch a job, to run right away, after a delay, or on a schedule:
svc := h.Service()
// Run as soon as possible
// Idempotency keys are scoped to the actor: dispatching again with the same key (for example, when retrying a request) doesn't enqueue a second job
_, _, err := svc.Dispatch(ctx, "user", user.ID, "send-welcome-email",
welcomeEmail{Name: user.Name},
actor.WithIdempotencyKey("welcome"),
)
// Run every Monday at 9am
_, _, err = svc.Dispatch(ctx, "user", user.ID, "send-weekly-summary", nil,
actor.WithJobCron("0 9 * * 1"),
)
Long-running tasks can also be scheduled across a cluster with the built-in task pool.
Workflows
Workflows are durable, multi-step processes. In Francis, a workflow is a graph of named steps, which can run in sequence or in parallel. Each step is implemented by a plain Go function, and it can have a compensation: a function that undoes it if a later step fails. Under the hood, a workflow runs as a built-in actor, so it survives restarts and the loss of a host, and its steps are spread across the cluster.
checkout, err := workflow.New("checkout",
workflow.WithSteps(
workflow.Step("reserve-inventory",
workflow.WithRun(reserveInventory),
workflow.WithCompensate(releaseInventory),
),
workflow.Step("charge",
workflow.WithRun(chargeCard),
workflow.WithCompensate(refundCharge),
workflow.WithMaxAttempts(3),
),
workflow.Step("create-shipment", workflow.WithRun(createShipment)),
),
)
if err != nil {
return err
}
// Register the workflow before starting the host, on every host that should run its steps
err = h.RegisterBuiltInActor(checkout)
If create-shipment fails, Francis runs refundCharge and then releaseInventory, and the workflow ends as failed.
Each step receives the workflow’s input (and, if it needs them, the outputs of earlier steps) and returns its own output:
func chargeCard(ctx context.Context, t workflow.Task) (any, error) {
var order Order
err := t.DecodeInput(&order)
if err != nil {
return nil, errors.Join(actor.ErrJobPermanentFailure, err)
}
// Steps are delivered at least once, so give the payment provider a stable idempotency key
idem := t.InstanceID() + "|" + t.Step()
charge, err := payments.Charge(ctx, order.PaymentMethod, order.Total, idem)
if err != nil {
return nil, err
}
return chargeResult{ChargeID: charge.ID}, nil
}
// Compensation for chargeCard
// Receives the step's output, so it knows which charge to refund
func refundCharge(ctx context.Context, c workflow.Compensation) error {
var res chargeResult
err := c.DecodeResult(&res)
if err != nil {
return errors.Join(actor.ErrJobPermanentFailure, err)
}
return payments.Refund(ctx, res.ChargeID, c.Cause())
}
Start returns as soon as the workflow is scheduled and the workflow runs asynchronously in background in the cluster:
svc := checkout.Service(h.Service())
// Using the order ID as the instance ID means that a retried request finds the run already in progress
id, _, err := svc.Start(ctx, order, workflow.WithInstanceID("order-"+order.ID))
if err != nil {
return err
}
status, err := svc.GetStatus(ctx, id)
Built-in actors
Francis also ships with built-in actors for common patterns: cron jobs that run once across the cluster, rate limiting, the task pool I mentioned above, and signals that release any number of waiting callers at once. They’re all things you’d otherwise end up building yourself, and they keep working when you add replicas.
For example, Pocket ID uses the rate limiter for its HTTP endpoints:
limiter, err := ratelimit.New("api",
// 100 requests per second for each key, with bursts of up to 20
ratelimit.WithRate(100),
ratelimit.WithBurst(20),
)
if err != nil {
return err
}
err = h.RegisterBuiltInActor(limiter)
// Then, in your HTTP middleware
rl := limiter.Service(h.Service())
allowed, retryAfter, err := rl.Allow(ctx, clientIP)
Each key (in this case a client IP) is its own actor, active on a single host at a time, so the limit is enforced across the cluster. The actor maintains the data in-memory, so performance should be comparable to using Redis (which is commonly used for rate-limiting in distributed systems, but requires running a separate database).
Getting started
Francis is open source under the MIT license, and you can find the code on GitHub. The documentation has a quickstart that gets the counter above running in a few minutes, plus guides on writing actors, choosing a topology, deploying the runtime, and building workflows.
If you give it a try, I’d love to hear how it goes. Feel free to open an issue on GitHub with any question, bug, or feedback!
