Skip to main content

DataStar.Tools

DataStar.Tools is the CLI that deploys DataStar packages against a target database. It's a self-contained .NET application that runs on Windows, Linux or macOS, so any build agent that can run .NET can run it.

What it does:

  • Deploys scripts from a manifest against Oracle or SQL Server.
  • Builds a deployable package straight from source control, with nothing checked out.
  • Optionally writes an audit trail (summary, history, and reversal tables) to the target database.
  • Optionally generates reversal (rollback) scripts, either into the audit tables or as a NuGet package.
  • Optionally publishes reversal packages to a NuGet repository.
  • Inflates data documents to SQL, manages shared definitions, and merges documents for git.

Deployment artefacts​

A typical release uses two or three artefacts:

  • DataStar.Tools package. The CLI itself. Published on the releases page as a zip or a NuGet package, and as a .NET tool. Install it once on the agent, or (recommended) deploy it as a release artefact so the deployment pulls in the right version.
  • Component templates. Required if reversal is enabled. Package the templates as a NuGet package (under a contents folder is fine) so the CLI can regenerate reversal scripts.
  • Deployment package. Produced by the build. Contains the scripts to run, a manifest listing the run order, and a metadata.json with build information (typically a work-item reference and version).

Running the CLI​

On Windows, run DataStar.Tools.exe. On Linux, run dotnet DataStar.Tools.dll. Every flag can also be set as a property in appsettings.json, where the key is the PascalCase form of the long flag (so --log-level becomes LogLevel).

DataStar.Tools.exe [options]

Flags are case-sensitive. Multiple manifests can be processed in one run: see --work-item for the regular-expression form that extracts a work-item reference from each manifest filename. Each work item is processed in its own transaction, so a failure rolls back only the current work item, not earlier ones in the batch.

Install it as a dotnet tool​

As well as the zip and the tarball, DataStar.Tools publishes as a .NET tool called datastar, so an agent with the .NET SDK can install it without unpacking anything:

dotnet tool install --global DataStar.Tools --version 3.1.0
datastar --help

The packages are per platform, so the install downloads only the one for the agent it runs on. The examples below use datastar; if you are running the zip or the tarball, substitute DataStar.Tools.exe or dotnet DataStar.Tools.dll and everything else is the same.

The tool needs the .NET 10 runtime, which any agent with the .NET SDK already has. For an agent with no .NET at all, use the zip or the tarball, which carry their own runtime.

Install it from your own NuGet feed​

The tool is several packages: DataStar.Tools, which is what you install, and one package per platform, DataStar.Tools.win-x64, DataStar.Tools.linux-x64, DataStar.Tools.linux-arm64, DataStar.Tools.osx-x64 and DataStar.Tools.osx-arm64. DataStar.Tools holds no program. It names the platform package to fetch, and the install downloads that one. So to install from your own feed (Azure Artifacts, Artifactory, Nexus, ProGet, or a folder), the feed needs DataStar.Tools and the platform package for every kind of agent you run.

  1. Download the packages from the download page, or from nuget.org, and check their SHA-256 checksums.

  2. Push them to your feed:

    dotnet nuget push "DataStar.Tools*.nupkg" --source https://your-feed/v3/index.json --api-key <key>
  3. Install from it, either by naming the feed:

    dotnet tool install --global DataStar.Tools --version 3.1.0 --add-source https://your-feed/v3/index.json

    or by adding the feed to the agent's nuget.config and running the plain dotnet tool install --global DataStar.Tools.

If an install fails saying it cannot find a package such as DataStar.Tools.linux-x64, the platform package for that agent is missing from the feed.

Commands​

--command-name (-cn) picks what to do. It defaults to release, so existing pipelines need no change.

CommandWhat it doesNeeds
releaseDeploys a package against a target database. The default.A connection; a licence key when reversal is enabled
reversalGenerates a reversal package from existing audit data.A connection
buildReads a deployment file and source control, and produces a deployable package.A git repository; no database, no workspace, no licence
inflateTurns data documents into the SQL they stand for.Nothing: no database, no licence
definitionsLists, prunes, shares or inlines the definitions a workspace's documents point at.The workspace, for share
mergegit's merge driver for data documents and changesets.The workspace's templates

build, inflate, definitions and merge never open a database connection and never ask for --license-key.

build: source control in, a package out​

A deployment file already names every item and the revision it is pinned at, so there is nothing to check out. build reads each item as a blob at its own revision, inflates any data document or changeset to SQL, and writes the scripts and a manifest.mf into an output folder. Give it a package file name and it packs that folder too.

datastar --command-name build \
--git-directory . \
--build-file releases/TASK-123/TASK-123.xml \
--database Microsoft \
--work-branch "$BUILD_SOURCEBRANCHNAME" \
--output-directory artifacts/scripts \
--package-directory artifacts/package \
--package-file 'TASK-123.zip'
FlagPurpose
--git-directory (-gd)The repository to read blobs from. Defaults to the current directory.
--build-file (-bf)Required. The deployment file, read from disk. Fetch it first if it lives on a work item; build does not read it out of a commit or a tag.
--database (-db)Required. Which vendor to inflate for: Microsoft or Oracle. A document inflates only for the vendor it was captured from.
--at (-at)Optional. Read every item, and every shared definition, at this one revision (a commit, tag or branch) instead of each item's pinned one. A deployment file in Branch mode is always read at one revision: --at when given, otherwise the repository's HEAD, which in a pipeline is the commit the build was started for.
--work-item (-wi)Optional. The work item outright. Wins over --work-branch.
--work-branch (-wb)Optional. A branch name to take the work item from when --work-item is not given, using the built-in patterns or --task-pattern.
--task-pattern (-tp)Optional. A regular expression that finds the work item in the branch name, tried before the built-in patterns.
--output-directory (-od)Where the scripts and manifest.mf are written. Defaults to build under the repository. It must be empty, because everything in it is packaged: a build refuses a folder that already holds files.
--cleanOptional. Empty the output directory before writing to it, rather than refusing it. Use it on an agent that reuses its workspace between builds.
--includeOptional, repeatable. A file or folder to add to the package root as it is, such as a variables file for the deployment step.
--package-file (-pf)Optional. Pack the output folder into this file, written to --package-directory. ${TaskId} (or ${WorkItem}) is replaced with the work item. A .nupkg name produces a NuGet package; anything else a zip.
--package-directory (-pd)Where the package is written. Required with --package-file. Also accepted as --artifacts-directory, which is what GitSources.Tools called it. The long name only: -ad means --audit-database here.
--package-id, --package-version, --package-author, --package-summaryFor a .nupkg: the package's id and version (otherwise both are parsed out of the file name), and its metadata. ${Manifest} in the summary is replaced with the deployment file's name.
--package-uri, --package-api-key, --package-headerPush the package to a NuGet feed once it is written.

What the output holds, per item in the deployment file:

  • a .sql script (or any file that is not a data document) is copied as it is, to <category>/<file>;
  • a data document or a changeset is inflated to <category>/<name>.sql, reading its shared definition out of the repository at the item's own revision, so no definitions folder travels with the package;
  • manifest.mf lists them in deployment order.

Beside the scripts, at the root:

  • the deployment file the build read, copied byte for byte, as deployment.xml or deployment.json according to its format;
  • metadata.json, holding the work item and the package version, for example {"WorkItem": "CTAM-2586", "Version": "26.9.191917"}. The Octopus DataStar release step reads these as its variables file. The version is left out when no package is being made.

Anything else the deployment stage needs goes in with --include, given once per file or folder: a file lands at the package root under its own name, and a folder brings its contents, laid out as they are. An include that would replace a file the build wrote is refused, naming it. Write such files anywhere but the output directory, which must be empty.

A release that bundles several work items can hold a changeset each for one component, and a package holds one file per component. build combines them into one changeset before anything is inflated, the same way the client does when a deployment runs: rows on different keys are taken as they are, and a row both work items changed is taken as a chain, the first one's before and the last one's after. The combined changeset records them both, as TASK-123+TASK-124, so the script says where its rows came from.

build is all or nothing. Any item that cannot be read at its revision, cannot be inflated, or would write over another item's file fails the build with Built nothing: N of M item(s) could not be built and exit code 1, and each reason is logged above it. Three things to expect:

  • an item that pins no Version needs --at, otherwise '<path>' pins no version, and the build was given no revision to read at;
  • changesets that cannot be chained, because the second does not start where the first left the row, are refused with both work items named. Deploy them separately, or rebase one on the other;
  • a changeset in the deployment beside its component's own document or script is refused (is in this deployment both as its component and as changeset TASK-123; deploy one or the other), because the two contradict each other: one applies a work item's rows, the other the whole component.

A NuGet package built this way records the repository it came from, read from the repository being built, so a package found months later says where its scripts came from. Any user name, token or password the clone URL carries (a CI job token, a personal access token) is left out. A clone with no remote simply records none.

Each item is read out of git as the client reads the same file from disk: a leading UTF-8 byte order mark is dropped, so a script an editor saved with one packages the same way the client would export it.

Because the documents are inflated here, the package holds scripts and a manifest and needs no templates, no workspace and no database when it lands. That is what lets the deploy stage run on a plain agent, and lets you decide after the build which environment to send it to.

tip

Nothing about build needs a task tracker. The work item comes from the branch name, so fetching a deployment file from Jira or Azure Boards belongs in the pipeline, not in the tool.

An end-to-end pipeline​

The build stage reads the repository and publishes a package; the release stage, on any agent, deploys it. Nothing in between needs templates or a workspace.

# Build stage: repository in, package out.
dotnet tool install --global DataStar.Tools --version 3.1.0
datastar --command-name build \
--git-directory . \
--build-file releases/TASK-123/TASK-123.xml \
--database Microsoft \
--work-branch "$BUILD_SOURCEBRANCHNAME" \
--output-directory artifacts/scripts \
--package-directory artifacts/package \
--package-file 'MyProduct.${TaskId}.${Version}.nupkg' \
--package-id MyProduct.TASK-123 \
--package-version "1.0.$BUILD_BUILDID"
# Writes artifacts/package/MyProduct.TASK-123.1.0.42.nupkg (for build 42).
# Publish artifacts/package with your pipeline's own step.

# Release stage: deploy the package. The scripts are plain SQL by now.
datastar --command-name release \
--deploy-package MyProduct.TASK-123.1.0.42.nupkg \
--working-directory artifacts/package \
--database Microsoft \
--connection-string "$UAT_CONNECTION_STRING" \
--environment-name UAT \
--work-item TASK-123 \
--version-number 42

To deploy from the unpacked folder instead, point --working-directory at artifacts/scripts and leave out --deploy-package. Add the audit and reversal flags below as you would for any release; reversal still needs --template-directory.

This replaces GitSources.Tools for a repository that holds data documents. GitSources copies the .json as it is and cannot include the definitions folder a shared-definition document needs, so release would fail to read it; build inflates at build time and the question does not arise.

inflate: documents to SQL, offline​

Turns every data document under a directory into the SQL script it stands for, with no database and no licence. Use it to review what a document deploys as, or to get scripts out of a checkout without a workspace.

datastar --command-name inflate --working-directory ./components
datastar --command-name inflate --working-directory . --output-directory ./inflated --database Microsoft
FlagPurpose
--working-directory (-wd)The folder to search. Every .json below it that is a data document is inflated; folders whose name starts with . are skipped. Defaults to the current directory.
--database (-db)Optional. Defaults to the vendor each document was captured from, which is the only vendor it can inflate for offline.
--output-directory (-od)Optional. Scripts are written beside their documents unless this is given, in which case the folder tree is mirrored under it.
--template-directory (-td)Optional. Only needed for a document that carries no script surface (an early 3.1 preview extract). When the directory is inside a workspace, the workspace's template folders are used.

Each document is inflated as it was extracted. A template that has changed shape since is reported (Template '<file>' has changed since the document was captured) as a warning for a document that carries its surface, and as a failure for one that does not. A shared definition is found in the definitions folder above the document, so run it from a checkout that has one. The command exits 1 if any document failed and says which.

definitions: manage shared definitions​

For workspaces whose data templates share their definition, which is the default.

datastar --command-name definitions --operation list --working-directory .
datastar --command-name definitions --operation prune --working-directory .
datastar --command-name definitions --operation share --working-directory . --category reference-data
datastar --command-name definitions --operation inline --working-directory .
OperationWhat it does
listEvery definition under definitions, how many documents point at it, and which are unused. A document pointing at a definition that is not there is reported as an error. The default.
pruneRemoves definitions no document in this working tree points at. Refuses to remove anything while a document cannot be read, since it might point at one.
shareRewrites documents that carry their definition inline as documents that point at a shared one, writing the definition. Only documents whose template says (or defaults to) Definition="shared" are converted; others are left as they are and said so.
inlineThe reverse: rewrites documents that point at a definition as documents that carry it.
FlagPurpose
--working-directory (-wd)A folder in the workspace. The workspace root is found by walking up to .ds/workspace.json; without one, the folder itself is the root.
--operation (-op)One of the four above. Defaults to list.
--category (-cy)Optional. Limits share and inline to one component category.
--template-directory (-td)Optional. share reads each document's template to know whether it shares; the workspace supplies the templates when this is not given.

A converted document is read back and compared with what it was before the file is replaced, so a conversion never changes what a document says; one that would is left untouched and reported.

warning

prune reads only the working tree in front of it. A definition nothing here points at may still be the one another branch's documents point at, and removing it merges into everyone. The command says so when it removes anything. Prune on main once the branches that used a definition have merged.

merge: the git merge driver for documents​

Merges a data document or a changeset by row rather than by line, so two people changing different rows of one component do not conflict. Git merges text by line, and two branches that both add a row in the same place, or both touch a table's last row, conflict as text although their rows do not.

DataStar does not register the driver for you. Register it once per clone, and declare it for your component and changeset folders in .gitattributes (committed, so every clone gets it):

git config merge.datastar.name "DataStar data documents and changesets"
git config merge.datastar.driver "datastar --command-name merge --merge-base %O --merge-ours %A --merge-theirs %B --merge-path %P"
components/**/*.json merge=datastar
changesets/**/*.json merge=datastar

Use your workspace's Component Location and Changeset Location if they are not the defaults. Definitions never need merging: two branches that captured the same schema wrote the same file.

FlagPurpose
--merge-ours (-mo)Required. git's %A: the current branch's version, which the result is written over.
--merge-theirs (-mt)Required. git's %B: the other branch's version.
--merge-base (-mb)git's %O: the common ancestor. Empty when the file is new on both sides, which is handled.
--merge-path (-mp)git's %P: the file's path in the working tree. git runs the driver on temporary files at the top of the working tree, so this is how the driver finds the file's workspace, and with it the templates and the definitions folder.
--template-directory (-td)Optional. Where to find templates when the clone has no .ds/workspace.json; add it to the driver command line.

The driver needs the document's template, resolved by name from the workspace; Could not merge <file>: template '<name>' was not found means the clone holds no workspace, or its template folders do not contain that template.

What it does, per table and row: a row changed on one side is taken; a row changed the same way on both is taken once; a row changed differently on both is a conflict. The merged file is written in the standard layout, with the rows in your branch's order and the other side's new rows placed after the row it lists before them. Then:

  • exit 0 when every row merged, and git carries on;
  • exit 1 when rows conflict. Each is logged by table and key (both sides changed table COUNTRY row COUNTRY_CODE = GB differently; ours was kept. Base ..., ours ..., theirs ...), your branch's version is kept for that row so the file stays a valid document, and git marks the file as conflicted. There are no conflict markers to hunt for: open the file, settle the listed rows, and git add it;
  • exit 1, with the file left as it was, when no row-level merge is possible: both sides changed the header (template, variables or changeset details) or a table's columns differently, or the file is not a document. Resolve those by hand, usually by extracting the component again on the merged branch.

A changeset merges the same way, by the row each change is for, so two people's changesets for one work item combine.

Flag reference​

Target database​

SwitchLong nameTypeDescription
-db--databaseStringRequired for every command but inflate, where it defaults to the vendor each document was captured from. Target database vendor: Oracle or Microsoft.
-cs--connection-stringStringThe database connection string. Oracle uses the ODP.NET format; SQL Server uses the SqlClient format. Oracle users should also consider --tns-admin and --wallet-location.
-en--environment-nameStringName of the target environment. Used in logging and as a default when a package description is not supplied.
-ds--default-schemaStringDefault schema for the deployment. Ignored if the templates specify a schema.
-ta--tns-adminStringDirectory containing the Oracle TNS admin files. Lets the connection string use a TNS alias instead of host / port / service.
-wl--wallet-locationStringPath to an Oracle Wallet. Used for shared credentials, so passwords don't appear in connection strings.

SQL Server connection encryption​

DataStar v3 ships with Microsoft.Data.SqlClient 6.x, which changed the default for two connection-string keys compared to the older driver bundled with v2.x:

Keyv2.x defaultNew driver default
EncryptFalseMandatory
TrustServerCertificateFalseFalse

Connection strings that worked under v2.x can therefore fail the TLS handshake under v3, typically against SQL Server instances that present self-signed or otherwise untrusted certificates. To smooth the upgrade, DataStar.Tools applies the v2.x defaults when the connection string sets neither key:

  • If your connection string contains neither Encrypt= nor TrustServerCertificate=, DataStar.Tools appends Encrypt=False;TrustServerCertificate=True; so the connection behaves the same as under v2.x. A single Information-level log line records exactly which defaults were added.
  • If you set either key explicitly, your value is used unchanged.

To force the new strict-TLS behaviour explicitly:

Server=...;Database=...;Encrypt=Mandatory;TrustServerCertificate=False;

To trust a self-signed certificate while still encrypting the connection:

Server=...;Database=...;Encrypt=Mandatory;TrustServerCertificate=True;

To match v2.x exactly (no encryption), either let the defaults apply or set both explicitly:

Server=...;Database=...;Encrypt=False;TrustServerCertificate=True;

Both Encrypt and TrustServerCertificate flow through --connection-string; no separate flags are exposed for them. In the Octopus action template, set the values in the dsr-connectionString parameter and the helper passes them through unchanged.

When a connection fails, the underlying SqlClient exception (certificate trust, authentication, name resolution, etc.) is now logged directly so the cause is visible in the deployment output, rather than a generic "Failed to connect to the Microsoft database" line as in earlier 3.0.x releases.

Deployment mode and manifest​

SwitchLong nameTypeDescription
-cn--command-nameStringThe command to execute: release (the default), reversal, build, inflate, definitions or merge. See Commands.
-wd--working-directoryStringDirectory containing the deployment packages or manifests; for inflate and definitions, the folder to work in. Defaults to the current directory.
-wi--work-itemStringWork item or user-story reference. Required for all releases. For a batch of manifests, set this to a regex (prefixed with @) that extracts the work item from the filename. For example, "@[A-Z]{2,}\\d+" extracts US123456 from 001-US123456.mf. For build, optional: the branch name supplies it otherwise.
-vn--version-numberStringVersion number being released. Required for all releases.
-mr--manifest.regexpStringRegex for matching manifest files. Defaults to any file with a .mf suffix.
-dp--deploy-packageStringName of a NuGet package to deploy, if you want to deploy directly from the package rather than from unpacked files.

Build​

SwitchLong nameTypeDescription
-gd--git-directoryStringThe repository the items are read out of. Defaults to the current directory.
-bf--build-fileStringThe deployment file to build, read from disk.
-at--atStringRead every item and definition at this revision instead of the version each item was pinned at.
-wb--work-branchStringA branch name to take the work item from when --work-item is not given.
-tp--task-patternStringThe pattern that finds the work item in a branch name. Defaults to the built-in patterns.

Definitions and merge​

SwitchLong nameTypeDescription
-op--operationStringThe definitions operation: list (default), prune, share or inline.
-cy--categoryStringLimits definitions to one component category.
-mb--merge-baseStringFor merge: the common base version (git's %O).
-mo--merge-oursStringFor merge: the current version, which the merge is written over (git's %A).
-mt--merge-theirsStringFor merge: the other branch's version (git's %B).
-mp--merge-pathStringFor merge: the file's path in the working tree (git's %P).

Audit tables​

SwitchLong nameTypeDescription
-ae--audit-enabledFlagWrite deployment details to the audit tables (summary, history, reversal).
-ah--audit-historyStringName of the audit history table. Required when auditing is enabled.
-ar--audit-reversalStringName of the audit reversal table. Required when auditing is enabled.
-as--audit-summaryStringName of the audit summary table. Required when auditing is enabled.
-ao--audit-schemaStringSchema where the audit tables live.
-ad--audit-databaseStringDatabase holding the audit tables. SQL Server only, for when audit tables are in a different database from the target.
-ai--audit-initializeFlagCreate the audit tables automatically if they don't already exist.

Reversal and packaging​

SwitchLong nameTypeDescription
-re--reversal-enabledFlagGenerate reversal scripts alongside the deployment. Requires --template-directory.
-td--template-directoryStringDirectory containing the DataStar component templates. Required when --reversal-enabled is set; optional for inflate, definitions and merge, which otherwise take the workspace's.
-ir--inverse-reversalFlagRun reversal scripts in reverse order. Oracle's default; recommended for SQL Server when using v2 templates.
-id--audit-idStringIn reversal mode, identifies the audit history record to generate a package from.
-ri--reversal-idStringIn reversal mode, identifies the reversal record to generate a package from.
--artifacts-directoryStringAnother name for --package-directory, as GitSources.Tools called it. No short form: -ad is --audit-database.
-pd--package-directoryStringOutput directory for the generated package (a reversal package, or build's).
-pf--package-fileStringOutput filename for the package. Supports ${TaskId} (or ${WorkItem}) and ${Version} substitutions.
-pi--package-invariantFlagGenerate a reversal package even when no changes were applied. Off by default.
--cleanFlagFor build: empty the output directory first. Without it, a build refuses an output directory that is not empty.
--includeString, repeatableFor build: a file or folder added to the package root as it is. Give it once per path.
-pk--package-idStringPackage id for the generated NuGet package.
-pv--package-versionStringPackage version for the generated NuGet package.
-pa--package-authorStringPackage author for the generated NuGet package.
-ps--package-summaryStringPackage summary for the generated NuGet package.
-pu--package-uriStringURI to publish the package to.
-px--package-api-keyStringAPI key for the package repository.
-ph--package-headerStringHeader name for the API key, if the repository expects something other than the default.

Connection and runtime​

SwitchLong nameTypeDescription
-st--statement-timeoutIntegerPer-statement timeout in seconds. 0 means no timeout on SQL Server.
-rt--rollback-transactionFlagRoll back the deployment after running it. Useful for dry-run verification before a real release.

Logging and output​

SwitchLong nameTypeDescription
-lf--log-fileStringAlso write logs to this file, in addition to the console.
-ll--log-levelStringLog verbosity. For example, debug enables debug logging.
-od--output-directoryStringOutput directory for artefacts: reversal packages, build's scripts and manifest, inflate's scripts.

Licence​

SwitchLong nameTypeDescription
-lk--license-keyStringThe software licence key supplied by Absolute Technology. Required when --reversal-enabled is set. Not used by build, inflate, definitions or merge.

Flag value types​

TypeMeaning
FlagNo value expected after the switch.
StringA string argument follows the switch.
IntegerAn integer argument follows the switch.
Booleantrue or false follows the switch.