DataStar.Tools
DataStar.Tools is the CLI that deploys DataStar packages against a target database. It's a self-contained .NET application that runs on Windows, Linux or macOS, so any build agent that can run .NET can run it.
What it does:
- Deploys scripts from a manifest against Oracle or SQL Server.
- Builds a deployable package straight from source control, with nothing checked out.
- Optionally writes an audit trail (summary, history, and reversal tables) to the target database.
- Optionally generates reversal (rollback) scripts, either into the audit tables or as a NuGet package.
- Optionally publishes reversal packages to a NuGet repository.
- Inflates data documents to SQL, manages shared definitions, and merges documents for git.
Deployment artefacts
A typical release uses two or three artefacts:
- DataStar.Tools package. The CLI itself. Published on the releases page as a zip or a NuGet package, and as a .NET tool. Install it once on the agent, or (recommended) deploy it as a release artefact so the deployment pulls in the right version.
- Component templates. Required if reversal is enabled. Package the templates as a NuGet package (under a
contentsfolder is fine) so the CLI can regenerate reversal scripts. - Deployment package. Produced by the build. Contains the scripts to run, a manifest listing the run order, and a
metadata.jsonwith build information (typically a work-item reference and version).
Running the CLI
On Windows, run DataStar.Tools.exe. On Linux, run dotnet DataStar.Tools.dll. Every flag can also be set as a property in appsettings.json, where the key is the PascalCase form of the long flag (so --log-level becomes LogLevel).
DataStar.Tools.exe [options]
Flags are case-sensitive. Multiple manifests can be processed in one run: see --work-item for the regular-expression form that extracts a work-item reference from each manifest filename. Each work item is processed in its own transaction, so a failure rolls back only the current work item, not earlier ones in the batch.
Install it as a dotnet tool
As well as the zip and the tarball, DataStar.Tools publishes as a .NET tool called datastar, so an agent with the .NET SDK can install it without unpacking anything:
dotnet tool install --global DataStar.Tools --version 3.1.0
datastar --help
The packages are per platform, so the install downloads only the one for the agent it runs on. The examples below use datastar; if you are running the zip or the tarball, substitute DataStar.Tools.exe or dotnet DataStar.Tools.dll and everything else is the same.
The tool needs the .NET 10 runtime, which any agent with the .NET SDK already has. For an agent with no .NET at all, use the zip or the tarball, which carry their own runtime.
Install it from your own NuGet feed
The tool is several packages: DataStar.Tools, which is what you install, and one package per platform, DataStar.Tools.win-x64, DataStar.Tools.linux-x64, DataStar.Tools.linux-arm64, DataStar.Tools.osx-x64 and DataStar.Tools.osx-arm64. DataStar.Tools holds no program. It names the platform package to fetch, and the install downloads that one. So to install from your own feed (Azure Artifacts, Artifactory, Nexus, ProGet, or a folder), the feed needs DataStar.Tools and the platform package for every kind of agent you run.
-
Download the packages from the download page, or from nuget.org, and check their SHA-256 checksums.
-
Push them to your feed:
dotnet nuget push "DataStar.Tools*.nupkg" --source https://your-feed/v3/index.json --api-key <key> -
Install from it, either by naming the feed:
dotnet tool install --global DataStar.Tools --version 3.1.0 --add-source https://your-feed/v3/index.jsonor by adding the feed to the agent's
nuget.configand running the plaindotnet tool install --global DataStar.Tools.
If an install fails saying it cannot find a package such as DataStar.Tools.linux-x64, the platform package for that agent is missing from the feed.
Commands
--command-name (-cn) picks what to do. It defaults to release, so existing pipelines need no change.
| Command | What it does | Needs |
|---|---|---|
release | Deploys a package against a target database. The default. | A connection; a licence key when reversal is enabled |
reversal | Generates a reversal package from existing audit data. | A connection |
build | Reads a deployment file and source control, and produces a deployable package. | A git repository; no database, no workspace, no licence |
inflate | Turns data documents into the SQL they stand for. | Nothing: no database, no licence |
definitions | Lists, prunes, shares or inlines the definitions a workspace's documents point at. | The workspace, for share |
merge | git's merge driver for data documents and changesets. | The workspace's templates |
build, inflate, definitions and merge never open a database connection and never ask for --license-key.
build: source control in, a package out
A deployment file already names every item and the revision it is pinned at, so there is nothing to check out. build reads each item as a blob at its own revision, inflates any data document or changeset to SQL, and writes the scripts and a manifest.mf into an output folder. Give it a package file name and it packs that folder too.
datastar --command-name build \
--git-directory . \
--build-file releases/TASK-123/TASK-123.xml \
--database Microsoft \
--work-branch "$BUILD_SOURCEBRANCHNAME" \
--output-directory artifacts/scripts \
--package-directory artifacts/package \
--package-file 'TASK-123.zip'
| Flag | Purpose |
|---|---|
--git-directory (-gd) | The repository to read blobs from. Defaults to the current directory. |
--build-file (-bf) | Required. The deployment file, read from disk. Fetch it first if it lives on a work item; build does not read it out of a commit or a tag. |
--database (-db) | Required. Which vendor to inflate for: Microsoft or Oracle. A document inflates only for the vendor it was captured from. |
--at (-at) | Optional. Read every item, and every shared definition, at this one revision (a commit, tag or branch) instead of each item's pinned one. A deployment file in Branch mode is always read at one revision: --at when given, otherwise the repository's HEAD, which in a pipeline is the commit the build was started for. |
--work-item (-wi) | Optional. The work item outright. Wins over --work-branch. |
--work-branch (-wb) | Optional. A branch name to take the work item from when --work-item is not given, using the built-in patterns or --task-pattern. |
--task-pattern (-tp) | Optional. A regular expression that finds the work item in the branch name, tried before the built-in patterns. |
--output-directory (-od) | Where the scripts and manifest.mf are written. Defaults to build under the repository. It must be empty, because everything in it is packaged: a build refuses a folder that already holds files. |
--clean | Optional. Empty the output directory before writing to it, rather than refusing it. Use it on an agent that reuses its workspace between builds. |
--include | Optional, repeatable. A file or folder to add to the package root as it is, such as a variables file for the deployment step. |
--package-file (-pf) | Optional. Pack the output folder into this file, written to --package-directory. ${TaskId} (or ${WorkItem}) is replaced with the work item. A .nupkg name produces a NuGet package; anything else a zip. |
--package-directory (-pd) | Where the package is written. Required with --package-file. Also accepted as --artifacts-directory, which is what GitSources.Tools called it. The long name only: -ad means --audit-database here. |
--package-id, --package-version, --package-author, --package-summary | For a .nupkg: the package's id and version (otherwise both are parsed out of the file name), and its metadata. ${Manifest} in the summary is replaced with the deployment file's name. |
--package-uri, --package-api-key, --package-header | Push the package to a NuGet feed once it is written. |
What the output holds, per item in the deployment file:
- a
.sqlscript (or any file that is not a data document) is copied as it is, to<category>/<file>; - a data document or a changeset is inflated to
<category>/<name>.sql, reading its shared definition out of the repository at the item's own revision, so nodefinitionsfolder travels with the package; manifest.mflists them in deployment order.
Beside the scripts, at the root:
- the deployment file the build read, copied byte for byte, as
deployment.xmlordeployment.jsonaccording to its format; metadata.json, holding the work item and the package version, for example{"WorkItem": "CTAM-2586", "Version": "26.9.191917"}. The Octopus DataStar release step reads these as its variables file. The version is left out when no package is being made.
Anything else the deployment stage needs goes in with --include, given once per file or folder: a file lands at the package root under its own name, and a folder brings its contents, laid out as they are. An include that would replace a file the build wrote is refused, naming it. Write such files anywhere but the output directory, which must be empty.
A release that bundles several work items can hold a changeset each for one component, and a package holds one file per component. build combines them into one changeset before anything is inflated, the same way the client does when a deployment runs: rows on different keys are taken as they are, and a row both work items changed is taken as a chain, the first one's before and the last one's after. The combined changeset records them both, as TASK-123+TASK-124, so the script says where its rows came from.
build is all or nothing. Any item that cannot be read at its revision, cannot be inflated, or would write over another item's file fails the build with Built nothing: N of M item(s) could not be built and exit code 1, and each reason is logged above it. Three things to expect:
- an item that pins no
Versionneeds--at, otherwise'<path>' pins no version, and the build was given no revision to read at; - changesets that cannot be chained, because the second does not start where the first left the row, are refused with both work items named. Deploy them separately, or rebase one on the other;
- a changeset in the deployment beside its component's own document or script is refused (
is in this deployment both as its component and as changeset TASK-123; deploy one or the other), because the two contradict each other: one applies a work item's rows, the other the whole component.
A NuGet package built this way records the repository it came from, read from the repository being built, so a package found months later says where its scripts came from. Any user name, token or password the clone URL carries (a CI job token, a personal access token) is left out. A clone with no remote simply records none.
Each item is read out of git as the client reads the same file from disk: a leading UTF-8 byte order mark is dropped, so a script an editor saved with one packages the same way the client would export it.
Because the documents are inflated here, the package holds scripts and a manifest and needs no templates, no workspace and no database when it lands. That is what lets the deploy stage run on a plain agent, and lets you decide after the build which environment to send it to.
Nothing about build needs a task tracker. The work item comes from the branch name, so fetching a deployment file from Jira or Azure Boards belongs in the pipeline, not in the tool.
An end-to-end pipeline
The build stage reads the repository and publishes a package; the release stage, on any agent, deploys it. Nothing in between needs templates or a workspace.
# Build stage: repository in, package out.
dotnet tool install --global DataStar.Tools --version 3.1.0
datastar --command-name build \
--git-directory . \
--build-file releases/TASK-123/TASK-123.xml \
--database Microsoft \
--work-branch "$BUILD_SOURCEBRANCHNAME" \
--output-directory artifacts/scripts \
--package-directory artifacts/package \
--package-file 'MyProduct.${TaskId}.${Version}.nupkg' \
--package-id MyProduct.TASK-123 \
--package-version "1.0.$BUILD_BUILDID"
# Writes artifacts/package/MyProduct.TASK-123.1.0.42.nupkg (for build 42).
# Publish artifacts/package with your pipeline's own step.
# Release stage: deploy the package. The scripts are plain SQL by now.
datastar --command-name release \
--deploy-package MyProduct.TASK-123.1.0.42.nupkg \
--working-directory artifacts/package \
--database Microsoft \
--connection-string "$UAT_CONNECTION_STRING" \
--environment-name UAT \
--work-item TASK-123 \
--version-number 42
To deploy from the unpacked folder instead, point --working-directory at artifacts/scripts and leave out --deploy-package. Add the audit and reversal flags below as you would for any release; reversal still needs --template-directory.
This replaces GitSources.Tools for a repository that holds data documents. GitSources copies the .json as it is and cannot include the definitions folder a shared-definition document needs, so release would fail to read it; build inflates at build time and the question does not arise.
inflate: documents to SQL, offline
Turns every data document under a directory into the SQL script it stands for, with no database and no licence. Use it to review what a document deploys as, or to get scripts out of a checkout without a workspace.
datastar --command-name inflate --working-directory ./components
datastar --command-name inflate --working-directory . --output-directory ./inflated --database Microsoft
| Flag | Purpose |
|---|---|
--working-directory (-wd) | The folder to search. Every .json below it that is a data document is inflated; folders whose name starts with . are skipped. Defaults to the current directory. |
--database (-db) | Optional. Defaults to the vendor each document was captured from, which is the only vendor it can inflate for offline. |
--output-directory (-od) | Optional. Scripts are written beside their documents unless this is given, in which case the folder tree is mirrored under it. |
--template-directory (-td) | Optional. Only needed for a document that carries no script surface (an early 3.1 preview extract). When the directory is inside a workspace, the workspace's template folders are used. |
Each document is inflated as it was extracted. A template that has changed shape since is reported (Template '<file>' has changed since the document was captured) as a warning for a document that carries its surface, and as a failure for one that does not. A shared definition is found in the definitions folder above the document, so run it from a checkout that has one. The command exits 1 if any document failed and says which.
definitions: manage shared definitions
For workspaces whose data templates share their definition, which is the default.
datastar --command-name definitions --operation list --working-directory .
datastar --command-name definitions --operation prune --working-directory .
datastar --command-name definitions --operation share --working-directory . --category reference-data
datastar --command-name definitions --operation inline --working-directory .
| Operation | What it does |
|---|---|
list | Every definition under definitions, how many documents point at it, and which are unused. A document pointing at a definition that is not there is reported as an error. The default. |
prune | Removes definitions no document in this working tree points at. Refuses to remove anything while a document cannot be read, since it might point at one. |
share | Rewrites documents that carry their definition inline as documents that point at a shared one, writing the definition. Only documents whose template says (or defaults to) Definition="shared" are converted; others are left as they are and said so. |
inline | The reverse: rewrites documents that point at a definition as documents that carry it. |
| Flag | Purpose |
|---|---|
--working-directory (-wd) | A folder in the workspace. The workspace root is found by walking up to .ds/workspace.json; without one, the folder itself is the root. |
--operation (-op) | One of the four above. Defaults to list. |
--category (-cy) | Optional. Limits share and inline to one component category. |
--template-directory (-td) | Optional. share reads each document's template to know whether it shares; the workspace supplies the templates when this is not given. |
A converted document is read back and compared with what it was before the file is replaced, so a conversion never changes what a document says; one that would is left untouched and reported.
prune reads only the working tree in front of it. A definition nothing here points at may still be the one another branch's documents point at, and removing it merges into everyone. The command says so when it removes anything. Prune on main once the branches that used a definition have merged.
merge: the git merge driver for documents
Merges a data document or a changeset by row rather than by line, so two people changing different rows of one component do not conflict. Git merges text by line, and two branches that both add a row in the same place, or both touch a table's last row, conflict as text although their rows do not.
DataStar does not register the driver for you. Register it once per clone, and declare it for your component and changeset folders in .gitattributes (committed, so every clone gets it):
git config merge.datastar.name "DataStar data documents and changesets"
git config merge.datastar.driver "datastar --command-name merge --merge-base %O --merge-ours %A --merge-theirs %B --merge-path %P"
components/**/*.json merge=datastar
changesets/**/*.json merge=datastar
Use your workspace's Component Location and Changeset Location if they are not the defaults. Definitions never need merging: two branches that captured the same schema wrote the same file.
| Flag | Purpose |
|---|---|
--merge-ours (-mo) | Required. git's %A: the current branch's version, which the result is written over. |
--merge-theirs (-mt) | Required. git's %B: the other branch's version. |
--merge-base (-mb) | git's %O: the common ancestor. Empty when the file is new on both sides, which is handled. |
--merge-path (-mp) | git's %P: the file's path in the working tree. git runs the driver on temporary files at the top of the working tree, so this is how the driver finds the file's workspace, and with it the templates and the definitions folder. |
--template-directory (-td) | Optional. Where to find templates when the clone has no .ds/workspace.json; add it to the driver command line. |
The driver needs the document's template, resolved by name from the workspace; Could not merge <file>: template '<name>' was not found means the clone holds no workspace, or its template folders do not contain that template.
What it does, per table and row: a row changed on one side is taken; a row changed the same way on both is taken once; a row changed differently on both is a conflict. The merged file is written in the standard layout, with the rows in your branch's order and the other side's new rows placed after the row it lists before them. Then:
- exit 0 when every row merged, and git carries on;
- exit 1 when rows conflict. Each is logged by table and key (
both sides changed table COUNTRY row COUNTRY_CODE = GB differently; ours was kept. Base ..., ours ..., theirs ...), your branch's version is kept for that row so the file stays a valid document, and git marks the file as conflicted. There are no conflict markers to hunt for: open the file, settle the listed rows, andgit addit; - exit 1, with the file left as it was, when no row-level merge is possible: both sides changed the header (template, variables or changeset details) or a table's columns differently, or the file is not a document. Resolve those by hand, usually by extracting the component again on the merged branch.
A changeset merges the same way, by the row each change is for, so two people's changesets for one work item combine.
Flag reference
Target database
| Switch | Long name | Type | Description |
|---|---|---|---|
-db | --database | String | Required for every command but inflate, where it defaults to the vendor each document was captured from. Target database vendor: Oracle or Microsoft. |
-cs | --connection-string | String | The database connection string. Oracle uses the ODP.NET format; SQL Server uses the SqlClient format. Oracle users should also consider --tns-admin and --wallet-location. |
-en | --environment-name | String | Name of the target environment. Used in logging and as a default when a package description is not supplied. |
-ds | --default-schema | String | Default schema for the deployment. Ignored if the templates specify a schema. |
-ta | --tns-admin | String | Directory containing the Oracle TNS admin files. Lets the connection string use a TNS alias instead of host / port / service. |
-wl | --wallet-location | String | Path to an Oracle Wallet. Used for shared credentials, so passwords don't appear in connection strings. |
SQL Server connection encryption
DataStar v3 ships with Microsoft.Data.SqlClient 6.x, which changed the
default for two connection-string keys compared to the older driver bundled
with v2.x:
| Key | v2.x default | New driver default |
|---|---|---|
Encrypt | False | Mandatory |
TrustServerCertificate | False | False |
Connection strings that worked under v2.x can therefore fail the TLS handshake under v3, typically against SQL Server instances that present self-signed or otherwise untrusted certificates. To smooth the upgrade, DataStar.Tools applies the v2.x defaults when the connection string sets neither key:
- If your connection string contains neither
Encrypt=norTrustServerCertificate=, DataStar.Tools appendsEncrypt=False;TrustServerCertificate=True;so the connection behaves the same as under v2.x. A single Information-level log line records exactly which defaults were added. - If you set either key explicitly, your value is used unchanged.
To force the new strict-TLS behaviour explicitly:
Server=...;Database=...;Encrypt=Mandatory;TrustServerCertificate=False;
To trust a self-signed certificate while still encrypting the connection:
Server=...;Database=...;Encrypt=Mandatory;TrustServerCertificate=True;
To match v2.x exactly (no encryption), either let the defaults apply or set both explicitly:
Server=...;Database=...;Encrypt=False;TrustServerCertificate=True;
Both Encrypt and TrustServerCertificate flow through --connection-string;
no separate flags are exposed for them. In the Octopus action template, set
the values in the dsr-connectionString parameter and the helper passes them
through unchanged.
When a connection fails, the underlying SqlClient exception (certificate
trust, authentication, name resolution, etc.) is now logged directly so the
cause is visible in the deployment output, rather than a generic
"Failed to connect to the Microsoft database" line as in earlier 3.0.x
releases.
Deployment mode and manifest
| Switch | Long name | Type | Description |
|---|---|---|---|
-cn | --command-name | String | The command to execute: release (the default), reversal, build, inflate, definitions or merge. See Commands. |
-wd | --working-directory | String | Directory containing the deployment packages or manifests; for inflate and definitions, the folder to work in. Defaults to the current directory. |
-wi | --work-item | String | Work item or user-story reference. Required for all releases. For a batch of manifests, set this to a regex (prefixed with @) that extracts the work item from the filename. For example, "@[A-Z]{2,}\\d+" extracts US123456 from 001-US123456.mf. For build, optional: the branch name supplies it otherwise. |
-vn | --version-number | String | Version number being released. Required for all releases. |
-mr | --manifest.regexp | String | Regex for matching manifest files. Defaults to any file with a .mf suffix. |
-dp | --deploy-package | String | Name of a NuGet package to deploy, if you want to deploy directly from the package rather than from unpacked files. |
Build
| Switch | Long name | Type | Description |
|---|---|---|---|
-gd | --git-directory | String | The repository the items are read out of. Defaults to the current directory. |
-bf | --build-file | String | The deployment file to build, read from disk. |
-at | --at | String | Read every item and definition at this revision instead of the version each item was pinned at. |
-wb | --work-branch | String | A branch name to take the work item from when --work-item is not given. |
-tp | --task-pattern | String | The pattern that finds the work item in a branch name. Defaults to the built-in patterns. |
Definitions and merge
| Switch | Long name | Type | Description |
|---|---|---|---|
-op | --operation | String | The definitions operation: list (default), prune, share or inline. |
-cy | --category | String | Limits definitions to one component category. |
-mb | --merge-base | String | For merge: the common base version (git's %O). |
-mo | --merge-ours | String | For merge: the current version, which the merge is written over (git's %A). |
-mt | --merge-theirs | String | For merge: the other branch's version (git's %B). |
-mp | --merge-path | String | For merge: the file's path in the working tree (git's %P). |
Audit tables
| Switch | Long name | Type | Description |
|---|---|---|---|
-ae | --audit-enabled | Flag | Write deployment details to the audit tables (summary, history, reversal). |
-ah | --audit-history | String | Name of the audit history table. Required when auditing is enabled. |
-ar | --audit-reversal | String | Name of the audit reversal table. Required when auditing is enabled. |
-as | --audit-summary | String | Name of the audit summary table. Required when auditing is enabled. |
-ao | --audit-schema | String | Schema where the audit tables live. |
-ad | --audit-database | String | Database holding the audit tables. SQL Server only, for when audit tables are in a different database from the target. |
-ai | --audit-initialize | Flag | Create the audit tables automatically if they don't already exist. |
Reversal and packaging
| Switch | Long name | Type | Description |
|---|---|---|---|
-re | --reversal-enabled | Flag | Generate reversal scripts alongside the deployment. Requires --template-directory. |
-td | --template-directory | String | Directory containing the DataStar component templates. Required when --reversal-enabled is set; optional for inflate, definitions and merge, which otherwise take the workspace's. |
-ir | --inverse-reversal | Flag | Run reversal scripts in reverse order. Oracle's default; recommended for SQL Server when using v2 templates. |
-id | --audit-id | String | In reversal mode, identifies the audit history record to generate a package from. |
-ri | --reversal-id | String | In reversal mode, identifies the reversal record to generate a package from. |
--artifacts-directory | String | Another name for --package-directory, as GitSources.Tools called it. No short form: -ad is --audit-database. | |
-pd | --package-directory | String | Output directory for the generated package (a reversal package, or build's). |
-pf | --package-file | String | Output filename for the package. Supports ${TaskId} (or ${WorkItem}) and ${Version} substitutions. |
-pi | --package-invariant | Flag | Generate a reversal package even when no changes were applied. Off by default. |
--clean | Flag | For build: empty the output directory first. Without it, a build refuses an output directory that is not empty. | |
--include | String, repeatable | For build: a file or folder added to the package root as it is. Give it once per path. | |
-pk | --package-id | String | Package id for the generated NuGet package. |
-pv | --package-version | String | Package version for the generated NuGet package. |
-pa | --package-author | String | Package author for the generated NuGet package. |
-ps | --package-summary | String | Package summary for the generated NuGet package. |
-pu | --package-uri | String | URI to publish the package to. |
-px | --package-api-key | String | API key for the package repository. |
-ph | --package-header | String | Header name for the API key, if the repository expects something other than the default. |
Connection and runtime
| Switch | Long name | Type | Description |
|---|---|---|---|
-st | --statement-timeout | Integer | Per-statement timeout in seconds. 0 means no timeout on SQL Server. |
-rt | --rollback-transaction | Flag | Roll back the deployment after running it. Useful for dry-run verification before a real release. |
Logging and output
| Switch | Long name | Type | Description |
|---|---|---|---|
-lf | --log-file | String | Also write logs to this file, in addition to the console. |
-ll | --log-level | String | Log verbosity. For example, debug enables debug logging. |
-od | --output-directory | String | Output directory for artefacts: reversal packages, build's scripts and manifest, inflate's scripts. |
Licence
| Switch | Long name | Type | Description |
|---|---|---|---|
-lk | --license-key | String | The software licence key supplied by Absolute Technology. Required when --reversal-enabled is set. Not used by build, inflate, definitions or merge. |
Flag value types
| Type | Meaning |
|---|---|
| Flag | No value expected after the switch. |
| String | A string argument follows the switch. |
| Integer | An integer argument follows the switch. |
| Boolean | true or false follows the switch. |