This tool attempts to identify Progenitor-based API dependencies between Rust components and show information about them. The focus is on providing information to inform the online upgrade project.
Quick start
Listing APIs and their consumers
$ cargo xtask ls-apis apis
...
Nexus Internal API (client: nexus-client)
consumed by: dpd (dendrite/dpd)
consumed by: omicron-nexus (omicron/nexus)
consumed by: omicron-sled-agent (omicron/sled-agent)
consumed by: oximeter-collector (omicron/oximeter/collector)
consumed by: propolis-server (propolis/bin/propolis-server)
...To see what Cargo dependencies caused the tool to identify an API dependency, use --show-deps:
$ cargo xtask ls-apis apis --show-deps
...
Dendrite DPD (client: dpd-client)
consumed by: ddmd (maghemite/ddmd)
via path+file:///home/dap/omicron-merge/out/ls-apis/checkout/maghemite/ddmd#0.1.0
consumed by: mgd (maghemite/mgd)
via path+file:///home/dap/omicron-merge/out/ls-apis/checkout/maghemite/mg-lower#0.1.0
via path+file:///home/dap/omicron-merge/out/ls-apis/checkout/maghemite/mgd#0.1.0
consumed by: omicron-nexus (omicron/nexus)
via path+file:///home/dap/omicron-merge/nexus#omicron-nexus@0.1.0
consumed by: omicron-sled-agent (omicron/sled-agent)
via path+file:///home/dap/omicron-merge/sled-agent#omicron-sled-agent@0.1.0
consumed by: tfportd (dendrite/tfportd)
via path+file:///home/dap/omicron-merge/out/ls-apis/checkout/dendrite/tfportd#0.1.0
consumed by: wicketd (omicron/wicketd)
via path+file:///home/dap/omicron-merge/wicketd#0.1.0
...These paths are local to your system and will differ from system to system.
Listing servers and the APIs exposed and consumed by each one
$ cargo xtask ls-apis servers
...
omicron-sled-agent (omicron/sled-agent)
exposes: Sled Agent (client = sled-agent-client)
consumes: ddm-admin-client
consumes: dns-service-client
consumes: dpd-client
consumes: gateway-client
consumes: mg-admin-client
consumes: nexus-client
consumes: propolis-client
consumes: sled-agent-client
...You can similarly use --show-deps to see the Cargo dependency path that shows the API dependency.
Listing deployment units, their servers, and the APIs produced and consumed by each one
Deployment units are groups of components that are always deployed together. For example, all the components in the host OS global zone and switch zone represent one deployment unit.
$ cargo xtask ls-apis deployment-units
...
Crucible
crucible-agent (crucible/agent)
exposes: Crucible Agent (client = crucible-agent-client)
crucible-downstairs (crucible/downstairs)
exposes: Crucible Repair (client = repair-client)
consumes: repair-client
Crucible Pantry
crucible-pantry (crucible/pantry)
exposes: Crucible Pantry (client = crucible-pantry-client)
...Visualizing dependencies
The servers and deployment-units commands accept a --dot argument to print output in a format that dot(1) can process:
$ cargo xtask ls-apis deployment-units --output-format dot > deployment-units.dotYou can generate a PNG image of the graph like this:
$ dot -T png -o deployment-units.png deployment-units.dotChecking the upgrade DAG
As of this writing:
All Dropshot/Progenitor APIs in the system are identified with this tool (as far as we know).
Every API is annotated with metadata in
api-manifest.toml.The API metadata specifies whether versioning for the API is managed on the server side exclusively or using client-side versioning as well. We’ve made this choice for every existing API.
There are no cycles in the server-side-managed-API dependency graph.
This tool verifies these properties. If you add a new API or a new API dependency without adding valid metadata, or if that metadata breaks these constraints, you will have to fix this up before you can land the change to Omicron.
You can verify the versioning metadata using:
$ cargo xtask ls-apis checkYou might see any of the following errors:
missing field `versioned_how`: this probably means that you added an API to the metadata but neglected to specify theversioned_howfield. This must be eitherclientorserver. In general, preferserver. The tool should tell you if you try to useserverbut that can’t work. If in doubt, ask the update team.at least one API has unknown version strategy (see above): this probably means you added an API withversion_how = "unknown". (The specific API(s) with this problem are listed immediately above this error.) You need to instead decide if it should be client-side or server-side versioned. If in doubt, ask the update team.missing field `versioned_how_reason`: this probably means that you added an API withversioned_how = "client", but did not add aversioned_how_reason. This field is required. It’s a free-text field that explains why we couldn’t use server-side versioning for this API.graph of server-managed components has a cycle (includes node: "omicron-nexus"): this probably means that you added an API withversioned_how = "server", but that created a circular dependency with other server-side-versioned components. You’ll need to break the cycle by changing one of these components (probably your new one) to be client-managed instead.API identified by client package "nexus-client" (Nexus Internal API) is the "client" in a "non-dag" dependency rule, but its "versioned_how" is not "client"(the specific APIs may be different): this (unlikely) condition means just what it says: the API manifest file contains a dependency rule saying that some API is not part of the update DAG, but that API is not marked accordingly. See the documentation below on filter rules.
Details
This tool is aimed at helping answer these questions related to online upgrade:
What Dropshot/Progenitor-based API dependencies exist on all the software that ships on the Oxide rack?
Why does any particular component depend on some other component?
Is there a way to sequence upgrades of some API servers so that clients can always assume that the corresponding servers have been upgraded?
This tool combines three sources of information:
The developer-maintained
package-manifest.toml, which defines exactly what winds up on a deployed systemCargo/Rust package metadata (including package names and dependencies)
Developer-maintained metadata about APIs and their dependencies, located in ./api-manifest.toml
This tool basically works as follows:
It loads
package-manifest.tomland uses that to determine which Git commits are used to deploy components from all of the relevant Cargo workspaces.It loads and validates information about all of the relevant Cargo workspaces by running
cargo metadatausing manifests from the local Git clones.Using this information, it identifies all packages that look like Progenitor-based clients for Dropshot APIs: these are packages that (1) depend directly on
progenitoras a normal or build dependency, and (2) end in-client. (A few non-client packages depend on Progenitor, likeomicron-common. These are ignored using a hardcoded ignore list. Any other package that depends on Progenitor but does not end in-clientwill produce a warning.)Then, it loads and validates the developer-maintained metadata (
api-manifest.toml).Then, it applies whatever filter has been selected and prints out whatever information was asked for.
Deployment units
The upgrade DAG’s nodes are deployment units: sets of Rust packages that are always shipped and updated together. For example, Sled Agent, Propolis, and the switch-zone components all update as one unit (the host OS image), so they form a single deployment unit.
Deployment units and their packages are declared in api-manifest.toml:
[[deployment_units]]
name = "Host OS"
packages = [
"omicron-sled-agent",
"propolis-server",
# ...switch zone components...
]Each entry in packages is a top-level package: a Rust binary shipped in that deployment unit. ls-apis walks each top-level package’s Cargo dependencies to discover the APIs it produces and consumes.
Embedded components
Sometimes one binary contains a body of code whose API dependencies are best tracked separately from the rest of it. The motivating case is the Rack Setup Service (RSS): it lives in the sled-agent-rack-setup library crate, which is linked into the omicron-sled-agent binary, but it runs only once (at rack initialization) and has different semantics from the rest of the Sled Agent.
An embedded component is such a library, tracked as an API consumer in its own right:
[[deployment_units]]
name = "Host OS"
packages = [ "omicron-sled-agent", ... ]
embedded_components = [
{ name = "sled-agent-rack-setup", inside = "omicron-sled-agent", lifecycle = "rack-init" },
]inside names the top-level package (in the same deployment unit) that links the embedded component. When ls-apis walks that package’s dependencies, the embedded component’s own crate is skipped, so any API dependency reachable only through the embedded component is attributed to the embedded component instead of to the containing binary. (A dependency the binary also reaches by another path is attributed to both.) The embedded component is then walked separately, as its own consumer.
An embedded component is assumed not to host any Dropshot APIs of its own, so embedded components are not walked as API producers.
Lifecycle
A component’s lifecycle describes when its code runs:
| Lifecycle | Meaning |
|---|---|
| The component runs during normal operation. Its API dependencies are part of the upgrade DAG. This is the default. |
| The component runs only at rack initialization, when the whole rack is at a single consistent version. It therefore cannot experience version skew during an online upgrade, so its API dependencies are excluded from the upgrade DAG and from cycle checks. |
lifecycle may currently be set only on an embedded component; every top-level package is treated as steady-state.
The filtering is a little complicated but very important!
The purpose of filtering
Built-in filtering aims to solve a few different problems:
Many apparent dependencies identified through the above process are bogus. This usually happens because a package
Pdepends on a Progenitor client solely for access to its types (e.g., to define aFromimpl for its own types). In this case, a component usingPdoes not necessarily depend on the corresponding API. We want to ignore these bogus dependencies altogether. (If the component does depend on that API, it must have a different dependency on the Progenitor client package and that one will still cause this tool to identify the API dependency.)While exploring the dependency graph, we sometimes want to exclude some legitimate dependencies. Sometimes, a package
Pdepends on a Progenitor client, but only for a test program or some other thing that doesn’t actually get deployed withP. These are not bogus dependencies, but they’re not interesting for the purpose of online upgrade.To keep track of (and filter output based on) developer-maintained labels for each API dependency. More on this below.
Our broader goal is to construct a DAG whose nodes are deployment units and whose edges represent API dependencies between them. By doing that, we can define an update order that greatly simplifies any changes to these APIs because clients can always assume their dependencies are updated before them. We hope to do this by:
Starting with the complete directed graph of API dependencies discovered by this tool, ignoring bogus dependencies and dependencies from non-deployed components.
Removing one edge, meaning that we nominate that API as one where clients cannot assume their dependencies will be updated before them.
Checking if we still have cycles. If so, repeat.
How filters work
Example
Filter rules are defined in api-manifest.toml in the dependency_filter_rules block. Here’s an example:
[[dependency_filter_rules]]
ancestor = "nexus-types"
client = "gateway-client"
evaluation = "bogus"
note = """
nexus-types depends on gateway-client for defining some types.
"""Implied in this rule is that the Rust package nexus-types depends on the Rust package gateway-client, which is a client for the MGS API. Without this rule, the tool would identify any Rust component that depends on nexus-types as depending on the MGS API. This rule says: ignore any dependency on gateway-client that goes through nexus-types because it’s bogus: it’s not a real dependency because nexus-types doesn’t actually make requests to MGS. It just borrows some types.
Say we have a component called omicron-nexus that depends on nexus-types and gateway-client. For that component, this rule has no effect because there’s another Rust dependency path from omicron-nexus to gateway-client that doesn’t go through nexus-types, so the tool still knows it depends on the MGS API.
But if we had a component called oximeter-collector that depends on nexus-types but doesn’t depend on gateway-client through any other path, then this rule prevents the tool from falsely claiming that oximeter-collector depends on the MGS API.
Evaluations
Filter rules always represent a determination that a human has made about one or more dependencies found by the tool. The possible evaluations are:
| Evaluation | Meaning |
|---|---|
| No determination has been made. These are included by default. This is also the default evaluation for a dependency, if no filter rules match it. |
| Any matching dependency is a false positive. The dependency should be ignored altogether. |
| The matching dependency is for a program that is never deployed, like a test program, even though the package that it’s in does get deployed. These are ignored by default. |
| Any matching dependency has been flagged as "will not be part of the DAG used for online upgrade". This is primarily to help us keep track of the specific dependencies that we’ve looked at and made this determination for. These are currently ignored by default. |
| Any matching dependency has been flagged as "we want this to be part of the DAG used for online upgrade". |
In summary:
All dependencies start as
unknown.All the known false positives have been flagged as
bogus.All the known dependencies from non-deployed programs inside deployed packages have been flagged as
not-deployed.What remains is to evaluate the rest of the edges and determine if they’re going to be
dagornon-dag.
It is a runtime error for two filter rules to match any dependency chain. This makes the evaluation unambiguous. i.e., you can’t have one rule match a dependency chain and say it’s bogus while another says it’s dag.
Applying different filters at runtime
By default, this command shows dependencies that might be in the final graph. This includes those labeled dag and unknown. It excludes bogus, non-dag, and not-deployed dependencies.
You can select different subsets using the --filter option, which accepts:
include-non-dag: show non-bogus, non-not-deployeddependencies (i.e., all dependencies that do exist in the deployed system). This includesnon-dagdependencies.non-bogus: show everything except bogus dependenciesbogus: show only the bogus dependencies (useful for seeing all the false positives)all: show everything, even bogus dependencies