diff --git a/README.md b/README.md index e1230fbf9f..08413b4950 100644 --- a/README.md +++ b/README.md @@ -35,7 +35,7 @@ the `launchSettings.json` file of each instance. When started in setup mode, the ## Migrating from RavenDB to SQL Server or PostgreSQL -See [Migrating from RavenDB to SQL Server or PostgreSQL](docs/migration/ravendb-to-sql-migration-instructions.md). +- [Migrate from RavenDB to SQL Server or PostgreSQL](docs/migration/ravendb-to-sql-migration-instructions.md): what an operator can run today, and what is not built yet. ## Secrets diff --git a/docs/migration/migration-system-design-diagram.png b/docs/migration/migration-system-design-diagram.png index 15845976ae..85636355e7 100644 Binary files a/docs/migration/migration-system-design-diagram.png and b/docs/migration/migration-system-design-diagram.png differ diff --git a/docs/migration/ravendb-to-sql-migration-instructions.md b/docs/migration/ravendb-to-sql-migration-instructions.md index d229d03d15..bf7a425abc 100644 --- a/docs/migration/ravendb-to-sql-migration-instructions.md +++ b/docs/migration/ravendb-to-sql-migration-instructions.md @@ -1,53 +1,234 @@ -# Migrating from RavenDB to SQL Server or PostgreSQL - -This page covers reporting on the RavenDB source before you migrate. How the migration works is in the [migration overview](ravendb-to-sql-migration-overview.md) and the [system design diagram](migration-system-design-diagram.png). +# Migration Engine: Instructions > [!NOTE] -> The source report sends RavenDB only reads, but loading a database lets RavenDB's own expiration, its automatic deletion of documents past their retention date, run against it. If you are keeping the RavenDB database as a fallback, back it up before you run the report, as [Goals](ravendb-to-sql-migration-overview.md#goals) explains. +> This build does not copy data yet. Only `--migration-source-report` works, and ServiceControl does not start if `ServiceControl/Migration/Enabled` is `true`. + +This page tells you how to move the data of a ServiceControl error instance from RavenDB to SQL Server or PostgreSQL. ServiceControl copies the data itself, so you do not need a separate tool. [The migration overview](ravendb-to-sql-migration-overview.md) explains how the migration works and why. + +The migration does not move the audit instance or the monitoring instance. The audit instance has no migration path and the monitoring instance keeps its data in memory, so it has nothing to move. + +## How the migration runs + +ServiceControl copies the data it needs to open while it is closed. **This part is your outage**. Then it opens and copies the event log and the archived and resolved failed messages in the background. **The migration never writes to RavenDB**. + +> [!IMPORTANT] +> Until ServiceControl opens, you can go back to RavenDB and lose nothing. After it opens, you cannot go back. + +The migration copies the data in categories, and each category is Copying, Done, Failed or Abandoned. [The overview](ravendb-to-sql-migration-overview.md#when-something-goes-wrong-you-decide-what-happens-next) explains the categories and the states. ## Before you start -The source is a ServiceControl error instance on RavenDB. Keep its RavenDB settings in its configuration: the migration reads RavenDB through them, including after `PersistenceType` is switched to SQL Server or PostgreSQL. +Make sure that your setup is supported: + +- The RavenDB source is one of these: + - embedded on the host (Windows installation only) + - self-hosted on your network, or RavenDB Cloud +- The primary database and the throughput database of RavenDB are on the same server or cluster. +- The SQL target is SQL Server or PostgreSQL. +- A SQL Server target has Full-Text Search installed. Message search needs it, so `--setup` fails without it. The stock SQL Server container image does not have it. PostgreSQL needs nothing extra. +- The ServiceControl host can connect to RavenDB, to the SQL database, and to the message body storage at the same time. +- If ServiceControl runs in a container, RavenDB is an external server. The container image does not contain the RavenDB server, so it cannot open an embedded database. + +Get the data ready: + +- Resolve or archive as many unresolved failed messages as you can. They never age out, and ServiceControl stays closed until they are all copied. **A large backlog makes the outage long**. +- Run `--import-failed-errors` while the instance is still on RavenDB. A failed error import is a message that ServiceControl took off the error queue but failed to store. The migration copies these either way, but imported ones arrive as ordinary failed messages. + +Keep these settings as they are: + +- Keep `ServiceControl/RetryHistoryDepth` above 0. If it is 0 or less, the migration refuses to start. At that value, the first completed retry after the move deletes all the copied retry history. +- Do not change `ServiceControl/ErrorRetentionPeriod` or `ServiceControl/EventRetentionPeriod` until the migration is finished. The migration uses them to decide which rows are too old to copy. If one changes after the copy starts, ServiceControl refuses to start. + +## Keep the RavenDB settings + +The migration reads RavenDB through the RavenDB settings that the instance already has. You do not add new settings for the source. Keep these settings in the configuration of the instance, also after you change `ServiceControl/PersistenceType` to SQL Server or PostgreSQL. | Setting | Environment variable | What it is | | --- | --- | --- | -| `ServiceControl/RavenDB/ConnectionString` | `SERVICECONTROL_RAVENDB_CONNECTIONSTRING` | An external RavenDB server. Leave unset for an embedded database | -| `ServiceControl/DbPath` | `SERVICECONTROL_DBPATH` | The embedded database's data directory | -| `ServiceControl/RavenDB/DatabaseName` | `SERVICECONTROL_RAVENDB_DATABASENAME` | The primary database, `primary` by default | -| `LicensingComponent/RavenDB/ThroughputDatabaseName` | `LICENSINGCOMPONENT_RAVENDB_THROUGHPUTDATABASENAME` | The throughput database, `throughput` by default | -| `ServiceControl/RavenDB/ClientCertificatePath` or `ServiceControl/RavenDB/ClientCertificateBase64`, with `ServiceControl/RavenDB/ClientCertificatePassword` | `SERVICECONTROL_RAVENDB_CLIENTCERTIFICATEPATH` and so on | A secured external server's client certificate | -| `ServiceControl/ErrorRetentionPeriod` | `SERVICECONTROL_ERRORRETENTIONPERIOD` | Required. Don't change it during the move | +| `ServiceControl/RavenDB/ConnectionString` | `SERVICECONTROL_RAVENDB_CONNECTIONSTRING` | The address of an external RavenDB server. Leave it unset for an embedded database. | +| `ServiceControl/DbPath` | `SERVICECONTROL_DBPATH` | The data directory of the embedded database. | +| `ServiceControl/RavenDB/DatabaseName` | `SERVICECONTROL_RAVENDB_DATABASENAME` | The primary database. The default is `primary`. | +| `LicensingComponent/RavenDB/ThroughputDatabaseName` | `LICENSINGCOMPONENT_RAVENDB_THROUGHPUTDATABASENAME` | The throughput database. The default is `throughput`. | +| `ServiceControl/RavenDB/ClientCertificatePath` or `ServiceControl/RavenDB/ClientCertificateBase64`, with `ServiceControl/RavenDB/ClientCertificatePassword` | `SERVICECONTROL_RAVENDB_CLIENTCERTIFICATEPATH` and so on | The client certificate for a secured external server. | +| `ServiceControl/ErrorRetentionPeriod` | `SERVICECONTROL_ERRORRETENTIONPERIOD` | Required. ServiceControl does not start without it. | -## Report on the source +In an environment variable, you can leave out the `SERVICECONTROL_` prefix. You cannot leave out the `LICENSINGCOMPONENT_` prefix. -Run the instance's executable with `--migration-source-report`: +## Step 1: Upgrade ServiceControl and make the source report + +> [!IMPORTANT] +> +> - Upgrade the instance to the latest version in the usual way, and keep it on RavenDB. The version that copies the data must be the version that last ran against RavenDB. +> - Do not start ServiceControl again until step 6. If ServiceControl starts on the new database without the migration turned on, it writes rows to the database, and the migration then refuses that database. If this happens, start again from a new, empty database. + +Then make the source report. The report shows what ServiceControl can read from RavenDB, so you find a wrong setting early. Start the instance executable with the `--migration-source-report` argument: ```powershell -# Installed on Windows, from the instance's installation folder +# Windows installation. Start this from the installation folder of the instance. .\ServiceControl.exe --migration-source-report ``` ```shell -# Container, against an external RavenDB server +# Container, against an external RavenDB server. docker run --rm --env-file servicecontrol.env ghcr.io/particular/servicecontrol: --migration-source-report ``` -From source, build `src/ServiceControl` and run the same command from its output folder, as in [How to run/debug locally](../../README.md#how-to-rundebug-locally). +Where you run the report depends on the source: + +- External server: run the report while ServiceControl runs. The report only reads. +- Embedded database: stop the ServiceControl service first. Run the report, then start the service again. The report starts its own RavenDB process on the data directory, and it cannot do that while the instance holds the directory. + +The report prints: + +- The source persistence that it read. +- The RavenDB server version. +- Whether the source is embedded or external, and where it is. +- The name of each of the two databases, and the setting that each name came from. +- A row count for every collection in both databases. + +The report exits with 0 when it prints a report, and with 1 when it cannot. If the report fails, read [Errors from the RavenDB source](#errors-from-the-ravendb-source). + +## Step 2: Create an empty SQL database + +1. Create a new database on SQL Server or PostgreSQL. +2. In the configuration of the instance, set `ServiceControl/PersistenceType` to `SQLServer` or `PostgreSQL`. +3. Set `ServiceControl/Database/ConnectionString` to the new database. +4. Set `ServiceControl/MessageBody/StorageType` to `FileSystem`, `AzureBlob` or `S3`, and set the settings for that storage type. +5. Run `ServiceControl.exe --setup`. This creates the SQL schema, which includes the table where the migration saves its progress. + +## Step 3: Turn the migration on + +1. Set `ServiceControl/Migration/Enabled` to `true`. +2. If you want less optional data, set a shorter window for the optional categories. A window is how far back an optional category copies. Set it to `0` to turn the category off. + +| Setting | Default | What it does | +| --- | --- | --- | +| `ServiceControl/Migration/EventLogWindow` | The value of `ServiceControl/EventRetentionPeriod` | How far back the event log copy goes. | +| `ServiceControl/Migration/ArchivedAndResolvedFailedMessagesWindow` | The value of `ServiceControl/ErrorRetentionPeriod` | How far back the copy of archived and resolved messages goes. | + +By default, the migration copies all the optional data that RavenDB still holds (current retension period setting). A window longer than the retention period copies nothing extra. After a category starts copying, you cannot change its window. To stop it early, abandon it. + +## Step 4: Run the dry run + +The dry run reads RavenDB and tells you what the migration will do, including a range for how long ServiceControl will be closed. It does not write anything. Run it with the `--migration-dry-run` argument, in the same way as the source report. [The overview](ravendb-to-sql-migration-overview.md#the-dry-run) lists what it reports. + +Book the outage from the range in the report. The range is the shortest time the outage can take, not a promise. If the report says that a required category will end Failed and that a retry cannot fix it, plan to abandon that category after it fails. + +## Step 5: Back up RavenDB + +> [!IMPORTANT] +> Back up both RavenDB databases, primary and throughput, before you start the copy. Or turn off RavenDB expiration on them. + +RavenDB keeps deleting old failed messages and event log items during the migration and after it. Without a backup, the RavenDB database that you keep to go back to loses data every day. + +## Step 6: Start ServiceControl + +Start the ServiceControl service, or the container. ServiceControl then does these things in order: -The report prints the RavenDB server version, whether the source is embedded or external and where it is, both database names with the setting each came from, and a row count for every collection in both databases. +1. It runs every startup check. If a check fails, ServiceControl does not start and names the check. It copies nothing. +2. It copies every required category while it is closed. If a required category fails, ServiceControl stays closed. It says which category failed, why, and which command to run. +3. It opens when every required category is Done or Abandoned. You do not need to restart it. +4. It copies the optional categories in the background. -- **External server:** run it while ServiceControl is running. It sends only reads, and the note above about expiration applies to a server you are keeping as a fallback. -- **Embedded database:** stop the ServiceControl service, run the report, then start the service again. The report starts its own RavenDB process against the data directory, which cannot happen while the instance holds it. -- **Container with an embedded database:** not supported, because the container image does not ship the RavenDB server. Point the instance at an external RavenDB server instead. +> [!IMPORTANT] +> Do not start any `--error-ingestion-only` worker until ServiceControl opens. A worker refuses to start while a required category is unfinished, but a worker that starts at the same time as ServiceControl can start before the copy saves its first progress, and nothing stops it. -## If the report fails +If something goes wrong, read [If something goes wrong](#if-something-goes-wrong). -The error says what to fix: +## Step 7: Watch the background copy -- **"has no database named ..."**: the database name setting it quotes is wrong. -- **"refused its client certificate access ..."**: grant that certificate Read access to the database, or supply a certificate that has it. -- **"could not start a server for the embedded database ..."**: a ServiceControl instance is still running against that data directory and holds it. Stop the instance, run the report, then start it again. +ServiceControl is open and works as usual while the optional categories copy. You can see the progress in three places: -## Not available yet +- The ServicePulse activity feed. It shows when a category fails and when RavenDB cannot be reached. +- The ServiceControl log. +- The output of `--migration-status`. This command does not open RavenDB, so you can run it at any time. + +If the copy slows down your production work, raise `ServiceControl/Migration/ThrottlePauseMilliseconds` and restart. This is the pause between two background batches. The default is 100 milliseconds. + +> [!IMPORTANT] +> Turning the migration off does not stop the copy, because ServiceControl refuses to start until every category is Done or Abandoned. See [how to abandon a category](#a-category-is-failed). + +Keep RavenDB running until every category is Done or Abandoned. If RavenDB goes down after ServiceControl opens, ServiceControl keeps running. The optional categories wait for RavenDB to come back, or you can abandon them. + +## Step 8: Finish the migration + +1. Stop ServiceControl. +2. Run `--migration-verify`. It counts the rows in RavenDB again, and compares them with the rows that the copy read. It also shows the skipped and merged rows. It exits with 0 only when every category is Done or Abandoned. +3. Set `ServiceControl/Migration/Enabled` to `false`. +4. Start ServiceControl. + +Verify reads all the data in RavenDB again, but not the message bodies. On a large instance this takes time, and ServiceControl is stopped during that time. + +If verify shows rows that the copy did not read, keep RavenDB. Those rows are only in RavenDB. RavenDB can show fewer rows than the copy read when nothing is wrong, because RavenDB deletes old rows. [The overview](ravendb-to-sql-migration-overview.md#the-dry-run) explains why. + +If you turn the migration off before every category is Done or Abandoned, ServiceControl refuses to start. The refusal names each unfinished category and the ways to continue. + +## If something goes wrong + +### A startup check fails + +ServiceControl does not start, and its log names the check that failed. Nothing is copied. Correct the cause, then start ServiceControl again. The dry run runs the same checks, so you can use it to test your correction. + +### A category is Failed + +A category becomes Failed when it cannot copy rows, or when an error stops it. The rows that it already copied stay in SQL. A restart does not change a Failed category. You must run one of two commands. + +#### To retry a category + +1. Stop ServiceControl. +2. Read why the category failed. The reason is in the startup refusal and in the output of `--migration-status`. +3. Correct the cause. It is usually outside the migration, for example the body storage cannot be reached, a certificate expired, a database is down, or a disk is full. +4. Run `--migration-retry ` to reset the migration progress for that category. +5. Start ServiceControl. It copies the category again from the start. + +#### To abandon a category + +1. Stop ServiceControl. +2. Run `--migration-abandon `. When you abandon a category, the categories that depend on it are abandoned too. +3. Start ServiceControl. + +> [!IMPORTANT] +> **Abandon is final**. The rows that the category already copied stay in SQL, and the rest are not copied. Abandon a category when a retry cannot fix it, or when the data is not worth the time. You cannot retry the event log, so you can only abandon it. + +You can abandon a required category only after it starts copying. You can abandon an optional category at any time, for example to stop a large copy early. + +### You want to go back to RavenDB + +You can go back to RavenDB without losing data only while ServiceControl is still closed: + +1. Set `ServiceControl/Migration/Enabled` to `false`. +2. Set `ServiceControl/PersistenceType` back to RavenDB. +3. Start ServiceControl. + +You lose only the copy. To try the migration again later, start from step 2 with a new, empty SQL database. + +> [!IMPORTANT] +> Do not use the old copy again. The migration skips the categories that it finished, and misses everything that RavenDB received after that. + +After ServiceControl opens on SQL, you cannot go back. You can only finish the migration, or abandon what is left. + +### Errors from the RavenDB source + +The source report, the dry run and the migration show these errors when they cannot read RavenDB: + +- `has no database named ...`: the database name in the setting that the message names is wrong. Correct that setting. +- `refused its client certificate access ...`: give that certificate Read access to the database, or use a certificate that already has this access. +- `is secured but no client certificate is configured`: the connection string starts with `https://`. Set `ServiceControl/RavenDB/ClientCertificatePath` or `ServiceControl/RavenDB/ClientCertificateBase64`. +- `is valid from ... to ..., which does not include now`: the client certificate expired, or it is not valid yet. Use a current certificate. +- `ServiceControl expects RavenDB Server version ... or higher`: the external RavenDB server is older than this version of ServiceControl supports. Upgrade the RavenDB server. +- `could not start a server for the embedded database ...`: the rest of the line gives the reason. Usually a ServiceControl instance still holds the data directory. Stop that instance and try again. The other reasons are a port that is already in use, and a missing RavenDB server. +- `did not finish loading within 5 minutes`: another process holds the embedded data directory, or the directory is damaged. You cannot change the 5 minutes. + +## Commands and settings + +On an embedded source, stop ServiceControl before you run a command that opens RavenDB. **Only one RavenDB process can use the data directory**. The overview lists [every command and when it can run](ravendb-to-sql-migration-overview.md#the-dry-run), [the retry and abandon commands](ravendb-to-sql-migration-overview.md#retry-and-abandon), and [every migration setting](ravendb-to-sql-migration-overview.md#monitoring-and-settings). + +### Run a command in a container + +In a container, run each command as a one-off container of the same image, against the same databases: + +```shell +docker run --rm --env-file servicecontrol.env ghcr.io/particular/servicecontrol: --migration-status +``` -Copying the data (`MigrationMode`), the dry run, and the status and verify commands are planned but not built. The planned steps are in [Migration workflow](ravendb-to-sql-migration-overview.md#migration-workflow). +The RavenDB source must be an external server. diff --git a/docs/migration/ravendb-to-sql-migration-overview.md b/docs/migration/ravendb-to-sql-migration-overview.md index b1277bd1e8..613dca1680 100644 --- a/docs/migration/ravendb-to-sql-migration-overview.md +++ b/docs/migration/ravendb-to-sql-migration-overview.md @@ -1,403 +1,496 @@ -# Moving data from RavenDB to SQL +# Migration Engine: Design Overview -## Purpose +> [!NOTE] +> This build does not carry the whole migration yet. This page describes it as it will be when it ships. -The migration moves an error instance's data from RavenDB to SQL Server or PostgreSQL, so a customer can switch persisters and keep their data. +This page describes the migration as designed, not what is built today. [The instructions](ravendb-to-sql-migration-instructions.md) describe how an operator runs the migration, and the [system design](ravendb-to-sql-migration-system-design.md) provides the system architecture. See the [glossary](#glossary) for terms used throughout. -This covers the error instance only. The audit instance has no SQL persister, so a customer who finishes this migration still runs RavenDB for audit. +## Overview -## Strategy +### What this is -- Switch over first, and copy only what has to be copied. Retention does most of the work: error retention is between 5 and 45 days and event retention defaults to 14 days, so most of the source ages out on its own within weeks. That is why archived and resolved messages are optional rather than required. Retention would have deleted them anyway. -- The required set is what the instance needs the moment it opens. Most of it is small, but two parts grow: unresolved failures, and the last 7 days of event log, which on a busy instance holds a row for every failure, retry and submission. The strategy assumes you keep unresolved failures low by resolving and archiving. **A neglected instance breaks that assumption**: unresolved failures can legitimately be months old, and a large backlog of them makes the closed window long rather than short. The dry run is what tells you which case you are in. -- Anything not selected simply ages out of RavenDB, and the customer deletes the old database when they are ready. +- The migration moves an **error instance's** data from RavenDB to SQL Server or PostgreSQL, so a customer can switch persisters and keep their data. +- It runs inside ServiceControl itself. There is no separate tool. +- **The audit instance is out of scope.** There is no migration path for the audit instance. +- **The monitoring instance is out of scope.** It keeps its data in memory, so there is nothing to move. -## Goals +### TL;DR -- **Minimal downtime**. Only the required data copies with ServiceControl closed. Optional data copies in the background while it serves traffic. -- **All three RavenDB sources are supported**. Embedded, a container, or RavenDB Cloud, on one code path rather than three. -- **No writes through the client**. The copier never changes the source, but RavenDB's own expiration does: the primary database has it configured, and the sweep keeps deleting failed messages and event log items throughout the migration and for as long afterwards as the instance is left running. The old database is a fallback that degrades from the moment you start. -- **Abandonable up to a known point, and only up to that point**. While ServiceControl is closed the copy can be thrown away at no cost, because nothing but the copier has written to SQL and the migration has written nothing to RavenDB: see [the one point you can go back](#the-one-point-you-can-go-back). Once the host opens there is no way back at all. -- **No duplicates and no gaps**. Rows and the resume cursor, a marker of the last row copied, commit in one transaction, so a crash needs no reconciliation. -- **Every identifier anything depends on is carried across**. The event log, historic retry operations and pending integration events are renumbered, because nothing references their keys. -- **Refuse rather than half-migrate**. Every check runs before the first row moves, and a failure is a host that will not start. -- **No silent loss**. A migration cannot end with a selected category still in progress or halted, only with each one finished or explicitly abandoned. Abandoning is a deliberate choice, and an abandoned category lets the host open. Skipped rows are counted and reported. The one exception is rows RavenDB's own expiration deletes while the copy runs: they are never read, so they are not skips, and the recount is what tells them apart from a real loss. See [what does not come across](#what-does-not-come-across). -- **Bounded impact on a live instance**. The background copy waits a fixed pause between batches, and reads are streamed, so memory does not track the size of the database. -- **Known before it starts, visible while it runs**. A dry run reports what will move and how long ServiceControl is closed, and every category transition is reported as it happens. -- **Use existing functionality where possible**. Progress goes through custom checks and the activity feed, so no new client or screen is needed. - -## Non-goals - -- **Zero downtime.** The required data is copied with ServiceControl closed, so there is a real, if short, outage. -- **Reversible once ServiceControl opens.** Nothing copies SQL rows back to RavenDB, so once the host has served traffic there is no rollback of any kind. -- **Steerable while running.** No pause or resume, and no abort command. Going back during the closed window means stopping and reconfiguring, and changing anything else means editing configuration and restarting. -- **A general-purpose migration tool.** The source is always RavenDB and the target is always a ServiceControl EF Core persister, both at versions this build can read. -- **Custom migration UI via ServicePulse.** Custom checks and the event log report progress, and migration configuration and control are not available in the UI. - -## Supported migration scenarios - -The copier runs inside the ServiceControl host, so every row and every message body travels from the source, through ServiceControl, to the target. There is no database-to-database transfer, no backup and restore, and no replication. Whether a combination works therefore comes down to whether the ServiceControl host can reach both ends at once, and how long it takes comes down to how far the data has to travel. - -**Supported locations:** - -- **The RavenDB source**: embedded on the ServiceControl host (Windows installations only, see below), self-hosted on the same network in a container, VM or bare metal, or RavenDB Cloud -- **The SQL target**: SQL Server or PostgreSQL on the ServiceControl host, elsewhere on the same network, in a container, or as a managed cloud service such as Azure SQL, Amazon RDS or Google Cloud SQL -- **Any combination of the two**, subject to the requirements below - -**By where the data has to travel:** - -- **On-prem to on-prem.** The common case and the fastest. Embedded or self-hosted RavenDB to SQL on the same host or the same network. -- **On-prem to cloud.** RavenDB on the network, managed SQL in a cloud. Works, but each batch is one round trip, so write latency multiplies by the number of batches rather than being amortised away. -- **Cloud to on-prem.** RavenDB Cloud down to local SQL. Works, and the customer pays egress on everything copied, most of which is archived messages and their bodies. -- **Cloud to cloud.** Works, and is only sensible when ServiceControl runs alongside one of them. A host sitting on-prem between two clouds pulls every byte down and pushes it straight back up. -- **The message body store is a third location.** Bodies come out of RavenDB attachments and go wherever the target is configured to put them: a filesystem, Azure Blob or S3, with small text bodies kept inline in the database. That decision is made at the same time as the database move. - -**Infrastructure requirements:** - -- The ServiceControl host needs network access to the RavenDB source, the SQL target and the body store simultaneously -- Both RavenDB databases, primary and throughput, on one server or cluster -- A SQL Server target must have Full-Text Search installed. `--setup` checks `SERVERPROPERTY('IsFullTextInstalled')` and fails if it is absent, because message search is not optional. A stock SQL Server container image does not include it. PostgreSQL needs nothing extra, since its index is a GIN over `to_tsvector` -- A managed target's transient failures are survivable: retry on failure is on by default and there is no setting to turn it off - -**Not supported:** - -- Any server-to-server copy: no backup and restore, no RavenDB ETL or replication into SQL, no external data pipeline -- A host that can reach only one of the two databases at a time, so no staged move by way of an offline copy -- **An embedded RavenDB source when ServiceControl runs in a container.** Reading an embedded database means starting a RavenDB server process, and the container image does not carry one: `ServiceControl.Persistence.RavenDB.csproj:36` excludes the `RavenDBServer` directory from the artifact, and the copy that would restore it at `:44` is conditional on `CI` not being set, which the Dockerfile sets. A containerised instance migrating away from embedded RavenDB has to point at an external RavenDB server rather than at a data directory. Windows installations are unaffected: the installer unzips the server unconditionally -- Primary and throughput RavenDB databases in different locations -- Anything but RavenDB as the source, or anything but a ServiceControl EF Core persister as the target - -## Migration workflow - -1. Upgrade ServiceControl as normal, still on RavenDB. -2. Set four things in configuration: the new `PersistenceType`, its connection string, `MigrationMode=true`, and whether you want the one [optional](#optional) category, archived and resolved messages, copied. -3. Run `--setup` to create the SQL schema. It fails against a SQL Server instance without Full-Text Search installed. -4. Run the [dry run](#dry-run). It reports what it resolved as a source, what each category holds, and an estimate of how long ServiceControl will be closed. Read [what the dry run reports](#dry-run) before booking an outage around its estimate. -5. Start ServiceControl (`MigrationMode=true`). -6. Every check runs before a single row moves. If one fails the host does not start and names which, having copied nothing, so a wrong database name or unconfigured body storage costs a restart rather than a half-finished migration. -7. The copying of [required data](#required) starts, with ServiceControl still closed: the copy runs inside that same start, before the API begins listening and before any background service runs. How long it takes depends on how many unresolved failures you have and how busy the last 7 days were, and the [dry run](#dry-run) gives you an estimate. If a required category halts, the host stays closed until you fix the cause and restart, or abandon that category. -8. ServiceControl opens by itself the moment the required copy finishes, with no second restart to perform, and whatever [optional data](#optional) you asked for is copied in the background while the instance runs normally. You can watch it from ServicePulse custom checks and events, but not steer it. -9. You run the verification pass once the background job has completed, which reports row counts on both sides category by category, accounting for deliberate skips so a difference is explained rather than reported as a fault, then set `MigrationMode=false` and restart. Counts can differ in both directions without anything being wrong. SQL can hold more rows, because RavenDB keeps expiring rows the copier already took. SQL can also hold fewer, because once ServiceControl opens it sends pending integration events, removes group comments whose group has no failed messages left, and removes failed error imports once they are imported again. -10. RavenDB data can be removed. - -- If `MigrationMode=false` is set while a selected category is still incomplete, the startup is gated: it refuses and names exactly what is outstanding, or, where the [free abort](#the-one-point-you-can-go-back) is still open, starts with a warning that says so. See [turning migration mode off is a gated startup too](#turning-migration-mode-off-is-a-gated-startup-too). -- A category that ended *complete with errors* counts as complete and does not block, though its skipped count is printed so the loss is stated rather than silent. -- An explicit override exists for a customer who has changed their mind and accepts leaving data behind. It marks the outstanding categories as abandoned, which is a deliberate end state rather than a failure, so the progress check settles and the guard stays armed for any later migration. -- While a selected category is still unfinished, ServiceControl pauses its own clean-up: the retention sweep, the purge API, the heartbeat settings sync and throughput collection. Your SQL database grows until the copy finishes, and these restart on their own once it does. -- **Steps 5 to 7 are the abort window**, which is not the override above: see [the one point you can go back](#the-one-point-you-can-go-back). -- A category that stops because too many rows failed is *halted*, and it stays that way until someone acts: fix the cause and restart to carry on from where it stopped, or abandon it deliberately if you accept the loss. See [a halt stops one category, and clearing it is a restart](#a-halt-stops-one-category-and-clearing-it-is-a-restart). - -## Architecture +- **Retention does most of the work.** Error retention is 5 to 45 days and event retention defaults to 14 days, so most of RavenDB ages out on its own within weeks. +- **What the instance needs to open is copied while it is closed.** This is the *required* data. Most of it is small. +- **Everything else is copied in the background while it runs.** This is the *optional* data: the event log and archived and resolved messages. +- **Anything not copied ages out of RavenDB**, and the customer deletes the old database when ready. +- **The main thing that makes the outage long is a big backlog of unresolved failures.** They never age out, they are required, and they can legitimately be months old. The strategy assumes you keep them low by resolving and archiving. The *dry run* tells you which case you are in. ```mermaid -flowchart TB - cfg["Configuration + restart
the only way to change anything"] - checks["Custom checks + activity feed
progress, with no new client needed"] - - subgraph host["One ServiceControl host process, started with MigrationMode = true"] - direction LR - raven["RavenDB persister
own AssemblyLoadContext
read-only lifecycle"] - engine["MigrationEngine
categories, throttle,
dry run, verification"] - target["EF Core persister
SQL Server or PostgreSQL
own AssemblyLoadContext"] - raven -->|"IMigrationSource"| engine - engine -->|"IMigrationTarget"| target - end - - old[("Old RavenDB
read only, never written to")] - sql[("SQL Server or PostgreSQL
plus the checkpoint table")] - bodies[("Message body store
filesystem, Azure Blob or S3")] - - cfg --> host - host --> checks - old --> raven - target --> sql - target --> bodies +flowchart LR + A["Checks"] --> B["Copy required data
ServiceControl closed"] + B --> V["Required all Done
or Abandoned"] --> C["ServiceControl opens
optional data copies
in the background"] + C --> D["Everything Done or
Abandoned: turn
migration off"] ``` -- **The engine and the host know no store.** They deal in categories, cursors and counts. The source maps a category to what it reads and describes itself as labelled facts, the target maps a category to where it writes and how to count it, and each contributes its own startup checks. RavenDB to SQL is the only supported pair, and the engine does not depend on it. -- **The source reads the instance's own RavenDB settings**, so an existing customer sets nothing new. Leave them in place when switching `PersistenceType`. -- **Both persisters load into the same process**, each into its own `AssemblyLoadContext`. -- **The engine references neither assembly.** It knows only `IMigrationSource` and `IMigrationTarget`, and treats the resume cursor as an opaque value it passes from one to the other, so it can be tested against fakes on either side. +### When something goes wrong, you decide what happens next + +The end goal is simple: **either all the data comes across, or you explicitly give up on a named part of it.** -## Startup sequence +- **Each *category* of data is in one of four states:** Copying, Done, Failed or Abandoned. +- **A category goes Failed when something went wrong**, for example a row it could not copy. It stays Failed, and the status says why. +- **You act with one of two commands**, with ServiceControl stopped: + - `--migration-retry `, after fixing the cause: the next start copies the category again from the beginning. + - `--migration-abandon `: give up on what it has not copied, and keep what it has. +- **A restart never decides anything for you.** A Failed category stays Failed until you run a command, so a restart policy cannot loop or lose data. +- **ServiceControl opens once every required category is Done or Abandoned**, and the migration is finished once every category is. ```mermaid -flowchart TB - A["ServiceControl starts on SQL"] --> M{"MigrationMode?"} - - M -->|"On"| B["Open the SQL target
and the RavenDB source, read only"] - B --> D{"All checks pass?"} - D -->|"No"| E["Host does not start.
Names the failed check.
Nothing is copied."] - D -->|"Yes"| F["Copy the required categories.
The API is not listening."] - F -->|"A required category halts"| T["Host stays closed.
Fix the cause and restart,
or abandon the category."] - F -->|"Required categories settled"| G["ServiceControl opens.
New failed messages go to SQL."] - G --> H["Copy the selected optional categories
in the background"] - - M -->|"Off"| N{"Any category
outstanding?"} - N -->|"No"| L["ServiceControl opens.
RavenDB is not opened."] - N -->|"Yes"| O{"Override set?"} - O -->|"Yes"| P["Mark each outstanding category abandoned,
log what it leaves behind,
and open."] - O -->|"No"| Q{"Has this instance
ever opened on SQL?"} - Q -->|"No"| R["Open with a warning: point PersistenceType
back at RavenDB, or carry on
and lose the way back."] - Q -->|"Yes"| S["Host does not start.
Names every outstanding category,
its counts, and every route out."] +stateDiagram-v2 + [*] --> Copying + Copying --> Done + Copying --> Failed: something went wrong + Failed --> Copying: --migration-retry + Failed --> Abandoned: --migration-abandon + Copying --> Abandoned: --migration-abandon + Done --> [*] + Abandoned --> [*] ``` -**Checked before a single row moves:** +### Goals -- The SQL schema is current -- Message body storage is writable -- Both RavenDB databases are reachable -- The client certificate is valid, where the source is an external server -- The source is at a version this build can read -- The selected categories are valid -- `RetryHistoryDepth` is greater than zero. At zero or less, the first completed retry after the migration deletes the entire copied retry history, and no row count would ever show it +- **Short outage.** Only the required data copies while ServiceControl is closed. +- **Support every RavenDB hosting shape.** Embedded, self-hosted or RavenDB Cloud, on one code path. +- **Never write to RavenDB.** But RavenDB's own expiration keeps deleting old failed messages and event log items, during the migration and afterwards. So the old database is a fallback that degrades from the moment you start, which is why [the instructions](ravendb-to-sql-migration-instructions.md#step-5-back-up-ravendb) have you back it up first. +- **No duplicates and no gaps.** The copied rows and the *checkpoint* commit together, so a crash needs no clean-up. +- **Every identifier anything points at is kept.** Only the event log, historic retry operations and pending integration events get new ids, because nothing points at theirs. +- **Refuses rather than half-migrates.** Every check runs before the first row moves. +- **No silent loss.** Every skipped row is counted with a reason, and a migration cannot finish with a category unfinished unless someone deliberately *abandons* it. +- **Gentle on a live instance.** The copy is streamed, so memory does not grow with the database, and there is a pause between background batches. +- **Known before it starts, visible while it runs.** The dry run predicts the outage, and progress shows up as it happens. -### Turning migration mode off is a gated startup too +### Non-goals -*The right-hand branch above is the half a customer meets last and expects least, so it is worth reading before the migration starts rather than at the end of one.* +- **Zero downtime.** There is a real, if short, outage. +- **Rollback once ServiceControl opens.** Nothing copies SQL rows back to RavenDB. +- **Pausing or steering a running copy.** There is no pause, resume or abort. Decisions are two commands run with ServiceControl stopped, and every other change is a configuration edit and a restart. +- **A general-purpose migration tool.** The source is always RavenDB and the target is always one of ServiceControl's SQL persisters, both at versions this build can read. The migration engine was design to be expandable in the future. -Every startup on a SQL Server or PostgreSQL instance looks at the checkpoint table before ServiceControl opens, whether `MigrationMode` is on or off. That is what stops a migration ending by accident, and it costs nothing on an instance with nothing outstanding: one that has never migrated holds no checkpoint rows, and one whose categories all settled has none left open. Both start normally. A RavenDB instance never reaches the gate. +### The moving parts -With `MigrationMode` off and at least one category still outstanding, one of three things happens, and each is said out loud at startup rather than discovered weeks later: +```mermaid +flowchart LR + subgraph host["One ServiceControl process"] + src["RavenDB
source"] --> eng["Migration
engine"] --> tgt["SQL
target"] + end + old[("RavenDB
read only")] --> src + tgt --> sql[("SQL Server or PostgreSQL
plus the checkpoint table")] + tgt --> bod[("Body store
filesystem, Blob or S3")] +``` -- **The override is set.** Every outstanding category is recorded as abandoned, with its copied and skipped counts left as they are, and the host starts. Each one is logged saying what state it was in, how much it had copied, and that whatever it had not copied stays only in RavenDB. Abandoning is final: selecting that category in a later migration does not copy it again. -- **The override is not set, and this instance has never opened on SQL.** This is the [free abort](#the-one-point-you-can-go-back), so the host starts and warns rather than refusing. The warning names the two moves: stop now and point `PersistenceType` back at RavenDB, which discards the partial copy and costs nothing else, or carry on, which opens ServiceControl on a partly copied database and ends the free abort. It deliberately does not mention the override, because at that moment nothing is lost yet. -- **The override is not set, and this instance has already opened on SQL.** The host does not start. The error names every outstanding category, its state, its copied and skipped counts and its last error, and then the three routes out: restart with `MigrationMode=true` to let the copy finish or to resume a halted category once its cause is fixed, set the override to abandon what is outstanding and start without it, or, if RavenDB is already gone, abandon, because that is the only exit left. +- **Both persisters load into one ServiceControl process**, each in its own `AssemblyLoadContext`, and both are live during the copy. +- **The engine knows neither database.** It knows only `IMigrationSource` and `IMigrationTarget`, and deals in categories, *cursors* and counts. This allows it to be tested, and expanded in the future. + - The source maps each category to what it reads, and describes itself as labelled facts. + - The target maps each category to where it writes and how to count it. + - Each side adds its own startup checks. + - The engine passes the cursor from one side to the other without looking inside it. + - RavenDB to SQL is the only supported pair, but the engine does not depend on it. +- **The source reuses the instance's existing RavenDB settings**, so an existing customer configures nothing new for it. Leave them in place when switching `ServiceControl/PersistenceType`. -**Which is why the source stays until verification passes.** A customer who decommissions RavenDB while a category is outstanding has both doors shut: `MigrationMode=true` cannot start, because it opens the source before it copies anything, and `MigrationMode=false` refuses. Abandoning is then the only way to start the instance, and it is a real loss whose size is the counts in that message. +### A migration from start to finish -## Data to be migrated (Categories) +[The instructions](ravendb-to-sql-migration-instructions.md) define the migration steps. -### Required +- **The [free abort](#going-back-to-ravendb) window runs from the first start with migration on until ServiceControl opens.** Until then, you can go back to RavenDB and lose only the copy. +- **While required data is still copying, none of ServiceControl's own clean-up runs**: the retention sweep, the purge API, the heartbeat settings sync and throughput collection. The required copy runs before any of them starts, so they start only once it finishes. Optional categories never hold them back. -- Unresolved **and retry-issued** failed messages, with their bodies. Attempt history collapses to the newest attempt, because the SQL model has no attempts table. Retry-issued messages are required for the same reason unresolved ones are: issuing a retry deletes the expiry, so they never age out. Leaving one behind means the retry confirmation arrives with no row to mark resolved, and the message stays missing from the customer's list while the retry actually succeeded -- Message redirects -- Endpoint settings -- Known endpoints, including the monitored flag. One category, because the flag is a property of the endpoint row and cannot be copied without it -- Notification settings -- The licence trial end date -- Throughput history -- Retry operations, unacknowledged and historic. One category, because RavenDB holds both lists in a single document -- Licensing report masks -- The uploaded licensed endpoint details file, which nothing recomputes: skipping it means the customer re-downloads it from the licence portal and uploads it again -- Subscriptions -- The last 7 days of the event log, so the ServicePulse activity feed shows what led up to the failures you are about to act on the moment ServiceControl opens. The 7 days count back from whichever is earlier: the newest event in RavenDB, or the last time RavenDB stored one. Counting from the data rather than from when the copy runs means a restart copies the same window. The second limit is there because event times come from the endpoints, and an endpoint whose clock runs ahead would otherwise push the window forward. Older events are not copied and age out of RavenDB on their own: at the default 14-day event retention that is at most another 7 days of history -- Integration events still waiting to be sent when you switch over. There are only any if the old instance was falling behind or could not reach the broker, and each one is an event a subscriber has not yet received. They are sent once ServiceControl opens, later than they would have been. RavenDB already sends them in no particular order, so no ordering is lost -- Custom checks, with the status each last reported. They are required because not every check reports again: an endpoint that is down never does, and a check with no repeat interval reports only when its endpoint starts. Leaving one behind could hide a known failure until that endpoint restarts -- Failed error imports, with their bodies. Each is a failed message ServiceControl took off the error queue but could not ingest, so it exists nowhere else, and it never expires. After the move, the "Error Message Ingestion" custom check keeps flagging them and `--import-failed-errors` imports them into SQL. Importing them on RavenDB before you start is better still, and the [dry run](#dry-run) tells you how many there are -- Group comments, copied after the unresolved failed messages, so a note such as "do not retry this group" is there the moment ServiceControl opens. They follow those messages because a comment belongs to a failure group, and the groups ServicePulse shows first are built from unresolved failed messages. Every comment except a blank one is copied. A comment on a group whose messages are all archived or resolved waits while those messages copy, and once the migration settles ServiceControl's own clean-up removes any comment whose group has no failed messages left, exactly as it always does on SQL +## Design -### Optional +### Supported locations -- Archived and resolved failed messages: the biggest category by far, and most of the copying time +The copy goes from RavenDB, through the ServiceControl host, to SQL. There is no database-to-database transfer. RavenDB holds documents and SQL holds tables, so the rows have to be reshaped, and ServiceControl's own persisters already know how to read one and write the other. So what matters is whether the host can reach both ends, and how far the data travels. -### Not migrated +| | Supported locations | +| --- | --- | +| RavenDB source | Embedded on the host (Windows installations only), self-hosted on the network (container, VM or bare metal), or RavenDB Cloud | +| SQL target | SQL Server or PostgreSQL on the host, on the network, in a container, or managed (Azure SQL, Amazon RDS, Google Cloud SQL) | +| Body store | Wherever the target is configured to put bodies: filesystem, Azure Blob or S3. Small text bodies stay in the database. This choice is made at the same time as the database move | -- The RavenDB index definitions -- The transient in-flight collections, which are empty when nothing is running: `RetryBatches`, `RetryBatchNowForwardings`, `FailedMessageRetries`, `ArchiveOperations` and `UnarchiveOperations` -- `ArchiveBatches` and `UnarchiveBatches`, which exist only because of how RavenDB works -- The `ConnectedApplications` document, which only versions 6.0 and 6.1 wrote. Since 6.2 the MassTransit connector status that ServicePulse uses to turn features on and off comes from the connector's own heartbeat. ServiceControl holds that in memory and refills it when the connector next reports, so nothing reads the document -- Broker and audit service version details, which refill on the throughput collector's next run -- Failed message edit locks, which stop one failed message being edited twice. An edited message is resolved, so the lock only matters if the message fails again afterwards: on RavenDB it can then never be edited again, and after the move it can be edited once more -- Heartbeat state, which neither persister stores: ServiceControl rebuilds it in memory from live heartbeats after every restart. The list of known endpoints and which ones are monitored is copied, so after the move heartbeat monitoring behaves exactly as it does after any restart. An endpoint instance that is down sends no heartbeat, so it is counted as failing on the dashboard with no last heartbeat time, and no heartbeat alert is raised for it +| Route | What to expect | +| --- | --- | +| On-prem to on-prem | The common case, and the fastest: embedded or self-hosted RavenDB to SQL on the same host or network | +| On-prem to cloud | Works. Each batch is a round trip, so write latency adds up per batch | +| Cloud to on-prem | Works. The customer pays egress on everything copied, mostly archived messages and their bodies | +| Cloud to cloud | Works, but only sensible when ServiceControl runs next to one of them. A host on-prem between two clouds pulls every byte down and pushes it straight back up | -## What does not come across +**Requirements:** -**Whole categories are never copied.** Which ones, and why nothing needs them, is the [not migrated](#not-migrated) list above. Anything in an optional category you did not select is also never copied, and nothing later goes back for it. Neither is an event log item raised before the start of the 7-day window, which falls outside the [required](#required) event log window rather than being skipped. +- The host reaches the RavenDB source, the SQL target and the body store at the same time. +- The SQL database is new and empty: only the schema `--setup` created. +- Both RavenDB databases, primary and throughput, are on one server or cluster. The migration opens both through one `IDocumentStore`, and so does the licensing component. +- A SQL Server target has Full-Text Search installed, because message search needs it. + - The `AddFullTextSearch` EF Core migration checks `SERVERPROPERTY('IsFullTextInstalled')` and fails if it is absent. + - An EF Core migration runs once, so this is checked the first time `--setup` brings a database up to that migration, not on every `--setup`. + - A stock SQL Server container image does not include Full-Text Search. + - PostgreSQL needs nothing extra, because its search index is a GIN index over `to_tsvector`. +- Transient failures on a managed target are already survivable. Retry on failure is on by default, and no setting turns it off. -**Rows skipped one at a time, and counted.** Each of these shows up in the skipped count for its category, broken out by reason, so you can see how much went and why: +**Not supported:** -- A failed message whose `UniqueMessageId` is not a GUID. The target column is a `uniqueidentifier` and the value is never regenerated, because it is simultaneously the primary key, the ServicePulse URL, the retry correlation key and the body lookup key. -- A failed message with no processing attempts recorded against it. The SQL model keeps the newest attempt and derives the failure time, the failing endpoint and the exception from it, all of which are required columns, so a message with nothing to derive them from cannot be written at all rather than being written blank. -- A failed message whose body cannot be read after three attempts. **The whole message is skipped, not just its body**, because a message with no body is worse than no message. -- A subscription whose message type or transport address exceeds 200 characters. The target key columns are capped at 200 characters, so it cannot be stored at all. -- An archived or resolved failed message, or an event log item, already past its retention period. SQL's retention clean-up would delete it on its first pass, so it is counted rather than copied only to be deleted. -- Endpoint settings for an endpoint ServiceControl does not know. ServiceControl removes those settings shortly after it starts. -- A row missing a value SQL requires, such as a known endpoint with no name or host, or a failed message with no failing endpoint address. An empty group comment is left behind the same way: RavenDB can store one, but SQL never does. +- Any server-to-server copy: backup and restore, RavenDB ETL or replication, or an external pipeline. +- A host that can reach only one database at a time, so no staged move through an offline copy. +- An embedded RavenDB source when ServiceControl runs in a container. Reading an embedded database means starting a RavenDB server process, and the container image does not include one. +- Primary and throughput RavenDB databases in different places. +- Any source other than RavenDB, or any target other than a ServiceControl SQL persister. + +### What data moves + +There are eighteen *categories*: sixteen required and two optional. + +| Category | Kind | Why | +| --- | --- | --- | +| Unresolved and retry-issued failed messages, with bodies | Required | The point of the instance. Issuing a retry removes the expiry, so retry-issued messages never age out. Leave one behind and the retry confirmation finds no row to resolve, so the message stays missing even though the retry succeeded | +| Message redirects | Required | | +| Known endpoints, with the monitored flag | Required | The flag is part of the endpoint row | +| Endpoint settings | Required | Copied after known endpoints | +| Notification settings | Required | | +| Licence trial end date | Required | | +| Licensing endpoint records | Required | The per-endpoint rows in the throughput database | +| Throughput history | Required | Copied after licensing endpoint records, because each day's row hangs off one | +| Retry operations, unacknowledged and historic | Required | One category, because RavenDB holds both lists in one document | +| Licensing report masks | Required | | +| Uploaded licensed endpoint details file | Required | Nothing recomputes it. Losing it means downloading it again from the licence portal and uploading it again | +| Subscriptions | Required | | +| Integration events still waiting to be sent | Required | Each is an event a subscriber has not received yet. There are only any if the old instance fell behind or could not reach the broker. They are sent once ServiceControl opens, later than they would have been. RavenDB sends them in no particular order anyway, so no ordering is lost | +| Custom checks, with their last status | Required | Not every check reports again. An endpoint that is down never does, and a check with no repeat interval reports only when its endpoint starts. Leaving one out could hide a known failure until that endpoint restarts | +| Failed error imports, with bodies | Required | Messages taken off the error queue but not ingested. They exist nowhere else and never expire. After the move the "Error Message Ingestion" custom check keeps flagging them and `--import-failed-errors` imports them into SQL. Importing them on RavenDB first is better still | +| Group comments | Required | Copied after unresolved failed messages, so a note like "do not retry" is there when ServiceControl opens. Blank comments are not copied. A comment on a group whose messages are all archived or resolved is removed by ServiceControl's usual clean-up soon after it opens, because that group has no failed messages in SQL yet | +| Event log | Optional | The history behind the ServicePulse activity feed. On a busy instance it holds a row for every failure, retry and submission | +| Archived and resolved failed messages, with bodies | Optional | The biggest category by far, and most of the copying time | + +**Never copied, because nothing needs them:** + +- RavenDB index definitions. +- In-flight work collections, which are empty when nothing is running: `RetryBatches`, `RetryBatchNowForwardings`, `FailedMessageRetries`, `ArchiveOperations` and `UnarchiveOperations`. +- `ArchiveBatches` and `UnarchiveBatches`, which exist only because of how RavenDB works. +- The `ConnectedApplications` document, written only by versions 6.0 and 6.1. Since 6.2 the MassTransit connector status, which ServicePulse uses to turn features on and off, comes from the connector's own heartbeat. ServiceControl holds that in memory and refills it when the connector next reports, so nothing reads the document. +- Broker and audit service version details, which refill on the next throughput collection. +- Failed message edit locks, which stop one failed message being edited twice. An edited message is resolved, so the lock only matters if it fails again: on RavenDB it can then never be edited again, and after the move it can be edited once more. +- Heartbeat state, which neither persister stores and which is rebuilt in memory after any restart. The known endpoints and monitored flags are copied, so heartbeat monitoring after the move behaves exactly as after any restart. An endpoint that is down shows as failing with no last heartbeat time, and raises no alert. + +### Optional data windows + +- **Each optional category copies only rows inside its *window*:** events raised within it, or messages archived or resolved within it. +- **The window defaults to the instance's own retention period**: `ServiceControl/EventRetentionPeriod` for the event log, `ServiceControl/ErrorRetentionPeriod` for archived and resolved messages. So by default everything RavenDB still holds is copied. +- **`0` turns a category off before it starts.** A longer window than the retention period copies nothing extra, because older rows would only be deleted by SQL's clean-up. +- **The window counts back from when the category first started copying**, which the checkpoint records, so a restart copies the same window rather than one that has slid forward. +- **The window is fixed once the category starts.** A changed value, including a changed retention period behind a default window, is refused at startup, naming the category. To stop a started category short, abandon it. + +### Data that changes shape + +- **Attempt history collapses to the newest attempt.** SQL has no attempts table, so a message that failed five times arrives showing one attempt. This applies to every failed message, whatever its status. +- **Subscriptions that differ only in message-type version merge into one row**, because the SQL key leaves the version out. +- **On SQL Server, endpoint settings whose names differ only in case merge into one row.** The name column compares without case by default. PostgreSQL keeps both. + - The name column's own collation decides, not the database default. A case-sensitive database whose name column was given a case-insensitive collation still merges. + - The copy keeps the first in RavenDB document-id order. That id is a hash of the name, so which one survives is effectively arbitrary. + - The copy logs a warning naming both endpoints and the one it kept. + - The dry run lists each pair by asking SQL Server how the column compares. For unusual characters its list can differ from what the copy does. +- **Event log items, historic retry operations and pending integration events get new numbers.** Nothing refers to the old ones, so this is safe. + +The dry run counts both kinds of merge before anything moves. + +### Skipped rows + +Every skipped row is counted under a reason, and its id is written to the log. + +> [!IMPORTANT] +> The log is the only record of each skipped row. Look them up in RavenDB while you still have it. + +- *Fault* skips mean something went wrong. +- *Harmless* skips are rows SQL would have removed anyway. + +| Reason | Kind | Can a retry fix it? | +| --- | --- | --- | +| A message body that cannot be read after three tries, or that RavenDB does not hold at all. The whole message is skipped, because a message with no body is worse than none | Fault | Yes, once the cause is fixed | +| An error from the SQL target | Fault | Yes, once the cause is fixed | +| A `UniqueMessageId` that is not a GUID. It is the key, the ServicePulse URL and the retry key, so it is never regenerated | Fault | No | +| A failed message with no processing attempts, so nothing to fill its required failure time, failing endpoint and exception from | Fault | No | +| A subscription whose message type or address is over 200 characters, the SQL key limit | Fault | No | +| A row missing a value SQL requires, such as a known endpoint with no name | Fault | No | +| A blank group comment, which SQL never stores | Harmless | Not needed | +| An archived or resolved message, or event log item, already past its retention period | Harmless | Not needed | + +- **Any fault skip makes the category [Failed](#failed-categories)**, even one. Harmless skips leave it Done. +- **Endpoint settings for an endpoint ServiceControl does not know are copied, not skipped.** ServiceControl removes them shortly after it opens, by the same rule it always applies, so the migration does not repeat that rule. +- **Rows outside a window are not skips.** They are simply not part of the copy. +- **Rows RavenDB expires during the copy are not skips either.** A deleted row is never read, so it is an absence, not a loss. + - Expiration only deletes a document carrying `@expires`, and only two kinds ever get one: archived or resolved failed messages, and event log items (`ExpirationManager.cs`). + - A failed message loses its expiry whenever it becomes unresolved or retry-issued again: when it fails again, is unarchived, or is retried. + - The exception: a message that failed again after being archived or resolved on version 6.18 or earlier kept its old expiry, and upgrading does not remove it. So a few unresolved messages can still expire during the copy. + - Both shrinking categories copy in the background, so the copy has hours to race the sweep. A window shorter than the retention period keeps the copy clear of it, by the difference between the two. + +### Startup checks + +In this order, cheapest first: + +1. Source and target are the supported pair, RavenDB to SQL. Answered from settings alone. A RavenDB instance with migration on is refused here, because RavenDB is only ever the source. +2. The optional category windows are valid time spans of zero or more. +3. `ServiceControl/RetryHistoryDepth` is above zero. At zero or less, the first completed retry after the migration deletes the whole copied retry history, and no row count would ever show it. +4. The SQL schema is current. This also refuses a database upgraded without `--setup`. +5. Until a required category has started copying: the SQL database holds no ServiceControl data. `--setup` writes none, but even one plain start of ServiceControl writes rows. Copying into a database that already served on SQL would mix the two. +6. Message body storage is writable. The probe body it writes is deleted again. +7. The SQL target opens, and no optional category's window has changed since that category started. +8. The RavenDB source opens: + - an embedded server starts; + - on an external server, the client certificate is in date and an `https://` address has one, checked before the first request; + - an external server is at least the RavenDB client's version; + - both databases load, within 5 minutes each on an embedded source. + + If every required category is already Done or Abandoned, a source that will not open does not stop the start. ServiceControl opens, and the optional categories wait for the source. + +Checks stop at the first one that refuses, and the refusal names it. Once open, the source can add checks of its own. The RavenDB source adds none. -**Things that change shape, and are not counted as skips at all.** The dry run counts the two merges before anything moves. Attempt history and renumbering apply to every row of their kind, so there is nothing to count. They are also the ones to read twice: +```mermaid +flowchart LR + A["Checks pass?"] -->|No| X["Host stays closed
and says why"] + A -->|Yes| B["Copy every required
category"] + B --> C["All Done
or Abandoned?"] + C -->|No| X + C -->|Yes| D["ServiceControl opens"] + D --> E["Copy optional
categories"] +``` -- **Processing attempt history collapses to the newest attempt.** The SQL model has no attempts table. This affects every failed message that failed more than once, whether it is unresolved, archived or resolved. A message that failed five times arrives showing one attempt, and the other four are gone. -- **Subscriptions that differ only in message-type version merge onto one row**, because the target key carries the type name without the version. -- **Endpoint settings for two endpoint names that differ only in case merge onto one row on SQL Server**, because SQL Server's default collation compares names without case, so one of the two settings is kept. PostgreSQL keeps both, and so does a SQL Server database created with a case-sensitive collation. The dry run counts this one too, by asking SQL Server how the name column compares, though for unusual characters its count can differ from what the copy does. -- **Event log items, historic retry operations and pending integration events are renumbered.** Their keys are database identities and nothing references them, so this is safe, but the old numbers do not survive. +- **The host also stays closed** if a second instance is writing the same checkpoints. +- **When ServiceControl opens**, the target is marked as opened, which ends the [free abort](#going-back-to-ravendb). +- **Once it has opened, a RavenDB outage does not stop it.** Only the background copy needs the source then, so ServiceControl opens anyway, and the optional categories [wait for the source](#failed-categories) to come back. -**Rows RavenDB deletes while the copy is running are an absence, not a skip.** Expiration only deletes a document carrying `@expires`, and only two kinds ever get one: a resolved or archived failed message, and an event log item (`ExpirationManager.cs:34,41`). A failed message loses its expiry whenever it becomes unresolved or retry-issued again: when it fails again, when it is unarchived, or when a retry is issued. The exception is a message that failed again after being archived or resolved while the instance ran version 6.18 or earlier: it kept its old expiry, and upgrading does not remove it, so a few unresolved messages can still expire during the copy. Apart from those, only the archived and resolved messages category and the event log can shrink underneath the copier. Archived and resolved messages copy in the background, where the window is longest. The event log's 7 days copy while ServiceControl is closed, and at the default 14-day event retention even the oldest of them is a week from expiring on a source that stopped recently, so the sweep reaches the window only on a source left stopped for days before the move. A document the sweep removes before the stream reaches it is never read, so it is counted nowhere. The copier counts each category before it starts, and if the copy comes up short of that count it counts the source again: rows that no longer exist were removed by RavenDB and are an absence, while rows that still exist but were never read halt the category. The dry run's count is a snapshot rather than a promise, and the recount is what shows that RavenDB removed rows while the copy ran. +### Reading from RavenDB -**A category can finish with a small amount of loss and still count as complete.** A few skipped rows in a large table leave the category in a *complete with errors* state, which blocks nothing. Its skipped count is printed and the ids of the skipped rows are written to the log, so while the RavenDB database still exists you can go and look at exactly what did not make it. +- **The source opens RavenDB read-only**: connect, check the version, stop. It never runs RavenDB's database setup, and it refuses any write request. +- **Both source databases are opened through one `IDocumentStore`**, which is why they must be on the same server or cluster. +- **Upgrade on RavenDB first**, so the build doing the copy is the one that last ran against the source. Nothing checks this. A RavenDB start rewrites no stored document, so the copy reads exactly what this build's RavenDB persister would read. The only version check compares the RavenDB server to the RavenDB client, and only for an external source. +- **Distance to the source sets the pace.** The copier already holds each document from the stream, so each body costs one round trip rather than two. But it is one per message, and they are not batched. Egress out of RavenDB Cloud is billed to the customer. -## The one point you can go back +### Writing to SQL -While ServiceControl is closed and the required copy is running, nothing except the copier has written to SQL, and the migration has written nothing to RavenDB, which is still authoritative. RavenDB's own expiration still runs, though: unless you disabled it, it keeps deleting expired failed messages and event log items, as [Goals](#goals) describes. If you need your instance back, set `MigrationMode=false`, point `PersistenceType` back at RavenDB, and start. You lose the copy, not your data. To start again later, start from an empty SQL database: drop it, create it, and run `--setup` again. Do not reuse the old copy, because a second attempt skips every category the first one finished, so anything that reached RavenDB since is left behind, and integration events RavenDB has since sent would be sent again. +- **A failed message is written whole**, with its status unchanged. +- **`UniqueMessageId` keeps its value** but changes type, from a string in RavenDB to a `uniqueidentifier` column in SQL. It is the primary key, the ServicePulse URL, the retry correlation key and the body lookup key at once. +- **`StatusChangedAt` is rebuilt.** + - For archived and resolved messages it comes from `@expires`, the only place RavenDB records it. + - For unresolved and retry-issued messages it is the newest attempt's time, because the column is `NOT NULL` and cannot be left empty. That is harmless, because the retention sweep only considers archived and resolved rows. +- **Deciding whether a row is past retention needs two retention periods.** The source's turns `@expires` back into the status-change time. The target's current one decides whether that time is past the cutoff. +- **Bodies go through `IBodyStoragePersistence`**, which owns compression and the choice of filesystem, Azure Blob or S3. The copier applies the inline threshold itself, because that threshold lives on the ingestion path. It defaults to 102,400 bytes and is set by `ServiceControl/MaxBodySizeToStore`. +- **Throughput rows are written directly, not through the collector, and each day's count is set, not added.** + - Throughput is required, so it copies while ServiceControl is closed, before any collector has written to SQL. + - Setting makes the category safe to resume after a crash, where adding would double-count. The collector's own path does add (`LicensingDataStore.cs:198`). + - Copying the rows stops the audit and broker collectors gathering the same days again, because `LastCollectedDate` is worked out from the newest throughput row (`LicensingDataStore.cs:45`). + - The checkpoint stops a second pass overwriting days the collectors have written since. -That window closes the moment ServiceControl opens. From then on new failed messages are ingesting into SQL, RavenDB is no longer current, and there is no rollback: nothing copies SQL rows back. The choice at that point is to finish the migration or to accept losing whatever has not been copied. +### Batch size and throttling -## Reading from RavenDB +- **Batch size comes from the database.** SQL Server divides its parameter budget by the column count, PostgreSQL uses a flat 50 rows. +- **Background batches have a pause between them**, 100 ms by default, set by `ServiceControl/Migration/ThrottlePauseMilliseconds`. The first batch of each category is not paused. +- **Raise the pause if the copy competes with production.** Turning migration off is not a remedy: the host refuses to start until every category is Done or Abandoned. +- **The required copy is never paused**, because ServiceControl is closed and nothing competes with it. -- A dedicated read-only RavenDB lifecycle opens the source: connect, check the version, stop. It never calls `DatabaseSetup.Execute`. -- Both source databases must be on the same server or cluster (`LicensingDataStore.cs:35`). -- The source has to be at a ServiceControl version this build can read. ServiceControl stamps a version marker into the database on upgrade, because the RavenDB server version says nothing about which ServiceControl version wrote the data. A source without a marker, or one from a newer major version, is refused by name rather than misread. -- Duration scales with distance to the source. The copier already holds the document from the stream, so each body costs **one** round trip rather than two, but it is one per message and they are not batched. Egress out of RavenDB Cloud is billed to the customer. See [batching and throttling](#batching-and-throttling). +### Progress and checkpoints -## Writing to SQL +The checkpoint is one row per category, kept in the SQL database and created by `--setup`. It holds the category's *state*, its cursor, its copied, skipped and already-present counts, how many of the already-present rows were merges, the count per skip reason, how many rows SQL held for the category when it settled, timings, the last error, the window it copies with, and a version number to catch a second writer. -- A whole `FailedMessage` is written with its stored status intact. -- `UniqueMessageId` keeps its value, but converts type: the source holds a string and the target column is a `uniqueidentifier`. It is the primary key, the ServicePulse URL, the retry correlation key and the body lookup key at once. -- `StatusChangedAt` is reconstructed from `@expires` for resolved and archived messages, which is the only place RavenDB sets it. Unresolved and retry-issued messages normally have no `@expires` (see [the exception](#what-does-not-come-across) for messages from version 6.18 or earlier), so the copier uses the newest processing attempt's timestamp. The column is `NOT NULL`, so it cannot be left empty, but the value is harmless for those two: the retention sweep only considers resolved and archived rows, so an unresolved message never ages out whatever is written here. -- Message bodies go through `IBodyStoragePersistence`, which owns the compression threshold and the choice of filesystem, Azure Blob or S3. The copier applies the 102,400-byte inline threshold itself, because that threshold lives on the ingestion path rather than in `IBodyStoragePersistence`. -- Throughput rows are written directly rather than through the collector, and the write sets each day's count rather than adding to it. Throughput is a required category, so it copies while ServiceControl is closed, before any collector has written to SQL. Setting is what makes the category safe to resume after a crash, where adding would double-count. Copying the rows is also what stops the audit and broker collectors re-gathering the same days when the host opens, because `LastCollectedDate` is derived from the newest throughput row rather than stored (`LicensingDataStore.cs:45`). The checkpoint is what stops a second pass overwriting days the collectors have written since. -- Identifiers narrow on the way across, and the dry run counts every kind. What narrows, merges or cannot be stored at all is in [what does not come across](#what-does-not-come-across). +```mermaid +sequenceDiagram + participant E as Engine + participant S as RavenDB + participant T as SQL + loop each batch + E->>S: Read rows after the cursor, and their bodies + E->>T: Write rows + new counts + new cursor + Note over T: One transaction + E->>E: Stop as Failed if too much of this run was skipped + end + E->>T: Settle as Done or Failed +``` -## Batching and throttling +- **Progress never gets ahead of the data.** Rows and cursor commit together, so a restart never skips rows that were not written. +- **A crash costs only the batch in flight.** The next start carries on from the committed cursor. +- **A batch read twice is counted once.** Rows already in SQL come back as *already present*, never as copies or skips, so every category is safe to run twice. Unreadable bodies stay off the checkpoint until the write commits. +- **Every skip has a reason, or the save is refused.** The per-reason counts must add up exactly to the skipped count, and the target's counts must match what it committed. +- **The stored counts always describe the rows in SQL.** The target adds its own outcome and saves the result beside the rows, so nothing is saved later or separately. +- **A second writer is caught, not merged.** Each save carries the version it read, so two hosts, or a command and a host, pointed at one database cannot interleave. +- **A copy stopped between two categories still shows the ones it never reached.** Before copying anything, a start saves a *not started* row for each category it is about to copy, so every gate sees them as unfinished. +- **Only the copier and the two commands write the checkpoint.** The status and verify commands only read it. -- Batch size comes from the provider: SQL Server divides its own parameter budget by the column count, PostgreSQL uses a flat 50 rows. -- The throttle is a configurable pause between batches, defaulting to 100 ms. Raising it slows the background copy and eases the load on production. Turning `MigrationMode` off is not a remedy: while a selected category is unfinished the host refuses to start, unless you abandon that category. +### Failed categories -## Checkpointing and resume +A category is *Failed* when something went wrong. Everything already copied stays committed. Nothing is retried in the background and nothing waits for a timer: the category stays Failed, with the reason written on it, until you run a [command](#retry-and-abandon). -A copy that runs for hours will be interrupted at some point: a restart, a dropped connection, a machine reboot. The checkpoint is what makes an interruption cost only the batch that was in flight. It is one row per category, kept on the target and created by `--setup` along with the rest of the schema, and it is written in the same database transaction as the rows it describes. Only the copier writes to it; the status and verify commands read it. +**What makes a category Failed:** -**What one row holds:** the category it tracks, its state, the resume cursor, how many rows were copied, skipped and already present, a count per skip reason, how many rows the source held when the category started, when it started, when it last made progress, when it settled, the last error, and a version number used to spot a second writer. +- **Any fault skip by the end of the category**, even one. +- **Too many fault skips in this run, which stops it early.** More than 5% of the rows processed so far *and* more than 100 rows (`ServiceControl/Migration/HaltThresholdPercent`, `ServiceControl/Migration/HaltThresholdMinimum`). Needing both means a tiny category does not stop on one bad row, and a huge one does not stop on its 101st. Under the threshold a category runs to the end, so you get the whole list of bad rows in one pass. +- **The category's own counts not balancing:** every row read must be copied, skipped or already present. +- **In a required category only, an exception or a stall of 30 minutes with nothing committed**, such as the body store unreachable, a certificate expired, the source or target down, or the disk full. The error and the cursor are recorded. -**The states, and which ones a restart re-enters.** `Complete`, `CompleteWithErrors` and `Abandoned` are terminal, so a restart passes straight over the category. `Halted` and `Blocked` are not: a halt is resumed from its cursor once the cause is fixed, and a block clears itself once the category it waits on settles, which is how group comments end up behind unresolved failed messages. `NotStarted` and `InProgress` both mean there is work to do. +**What does not make it Failed:** -### One batch, and why nothing provisional is ever saved +- **Harmless skips.** +- **A shutdown or a crash.** The category stays Copying and the next start carries on from the cursor. +- **An exception or a stall in an optional category.** The error is written on the category, shown in status, and the next start carries on from the cursor. A RavenDB outage after opening is handled the same way. So a network blip never forces anything to be abandoned. If the error keeps coming back, retry or abandon the category. -```mermaid -sequenceDiagram - participant E as Migration engine - participant S as RavenDB source - participant T as SQL target - participant C as Checkpoint row - - E->>C: Read this category's row - C-->>E: State, cursor, totals so far - loop One batch at a time - E->>S: Read the next batch after the cursor - S-->>E: Rows, and the cursor they end at - opt The category carries message bodies - E->>S: Read each body, up to three attempts - S-->>E: The bodies, and which ones could not be read - end - E->>T: Write the rows, with the totals so far,
the unreadable bodies and the new cursor - Note over T,C: One transaction. The rows, the target's own skips,
the new totals and the cursor all commit, or none of them do - T-->>E: The checkpoint exactly as it committed - E->>E: Halt if too much of this run was skipped - end - E->>C: Settle as complete, complete with errors, or halted -``` +**What a Failed category costs:** -The thing to read twice is that the counts never travel back through the engine to be saved on some later write. The engine hands the target the totals so far, the target adds its own outcome to them and saves the result beside the rows, and the engine then keeps whatever committed. So there is no window in which the stored row claims rows that are not there, and a crash at any instant leaves counts and cursor that both describe exactly the rows in SQL. +- **The other categories still run.** Every required category gets its go in one start, so one fix-and-retry covers them all. A category that must follow a Failed one waits. +- **A Failed required category** keeps ServiceControl closed until it is retried to Done or abandoned. That is deliberate: opening the host is the point of no return. +- **A Failed optional category** leaves the instance running with that slice missing, and stops the migration being finished. -**What that buys, and why each part is needed:** +**What the checks cannot see, and accept:** -- **Progress never gets ahead of the data.** The rows and the cursor commit together, so a restart cannot skip past rows that were never written. -- **A crash costs the batch in flight and nothing else.** The next run reads from the committed cursor. -- **Re-reading a batch cannot double-count it.** Unreadable bodies stay off the checkpoint until the write commits, so a batch that is read twice is counted once, and rows the earlier attempt did write come back as *already present* rather than as fresh copies. -- **Every skipped row has a reason, or the save is refused.** The checkpoint rejects a batch reporting more skips than it explains, because verification has to account for each one rather than report a healthy migration as broken. -- **Each category resumes independently**, so a half-copied category picks up where it stopped while its neighbours are untouched. -- **The halt counters are per run and deliberately not stored.** If the skips that tripped a halt stayed on the row, a restart with the cause fixed would re-trip it on its first batch. -- **A second writer is caught rather than merged.** Each save carries the version it read, and a save against a row that has moved on is refused, so two hosts pointed at one target cannot quietly interleave their progress. -- **If a message is already in SQL the SQL row wins and the copier skips it**, which is what makes every category safe to run twice. +- **Nothing counts RavenDB during the copy.** So a read that ends early, or skips a range, still balances, and the category ends Done. `--migration-verify` finds it afterwards, by counting RavenDB again and comparing that with the rows the copy read (see [the read-only commands](#the-dry-run)). +- **Verify counts with the same reader the copy used.** So it finds rows one run missed, but not a reader that always misses the same rows, or a source pointed at the wrong database. Tests on each reader cover those. +- **A write that silently stores fewer rows still ends Done.** Already present is what is left over after copied and skipped, so the counts balance. The counts check does catch rows dropped while a batch is prepared. Verify prints how many already-present rows were merges, so on a category copied once, already-present rows that merges do not explain point to it. +- **The threshold judges this run only.** A bad start can stop a category whose overall rate would have been fine, and the source must not read the rows most likely to be skipped first. -## Error handling +### Retry and abandon -- Which rows are skipped, and why, is in [what does not come across](#what-does-not-come-across). What follows is the mechanics around those rules. -- A body is read up to three times before the message is skipped whole, and the exhausted attempts count toward the halt threshold. -- Deciding whether a row is past the target's retention cutoff needs two retention periods: the source's reverses `@expires` back into the status-change instant, and the target's current one decides whether that instant is past the cutoff. -- A bad row does not stop the copy. Its category finishes in a separate complete-with-errors state. -- The halt threshold is proportional, 5 percent by default, with an absolute floor, and skips halt a category only when both are exceeded. The one exception is a category smaller than the floor, which halts if it loses more than half its rows. Proportional alone halts a three-row category on one bad row; absolute alone halts a five-million-row table on its 101st failure at the default floor of 100. Together, a large category keeps going through losses under the percentage and finishes complete with errors, so ten thousand skipped rows out of five million do not halt it. -- Rows left behind because SQL would remove them anyway (past retention, settings for unknown endpoints) are counted and reported, but never halt a category. The target reports them apart from its real failures, so they land in the skipped count and the log without moving the category toward a halt. -- The percentage is measured against what the run has processed so far rather than against the category's total, so a run that starts badly looks worse than it is. The floor is what keeps that harmless, since fewer than 101 skipped rows never consults the percentage at all. More than that, bunched at the start, does halt a category whose overall rate would have been fine, and the cost is one restart: the skipped rows commit with the cursor, so the next run resumes past them with its counters back at zero. -- A source therefore must not read a category in an order that puts the rows most likely to be skipped at the front of it. -- Verification therefore cannot treat any count difference as a fault. It accounts for every skip rule, or it reports every successful migration as broken. +Stop ServiceControl, run one of these, then start it again. In a container, run the command as a one-off `docker run --rm ` against the same database, the same way as `--migration-source-report`. -### A halt stops one category, and clearing it is a restart +| Command | Works on | What it does | +| --- | --- | --- | +| `--migration-retry ` | A Failed category, or an optional one still copying with an error | The next start copies it again from the beginning | +| `--migration-abandon ` | A Failed or started required category, or any optional one | Gives up on what it has not copied, and on any category whose rows hang off it. What they copied stays in SQL. Final | -A halt is the copy refusing to keep going on one category because something is wrong beyond the odd bad row. It is not a crash and not data loss: everything already copied is committed, the cursor points at the row after the last one that committed, and the reason is written on the category. Nothing is retried in the background and nothing waits for a timer. The category sits halted until a person does something about it. +**To retry:** -```mermaid -stateDiagram-v2 - [*] --> NotStarted: nothing has run yet - NotStarted --> InProgress: the host starts with MigrationMode = true - NotStarted --> Blocked: the category it must follow has not settled - Blocked --> InProgress: that category settles, then the next restart - InProgress --> InProgress: the host was stopped mid-copy,
so the next start resumes from the cursor - InProgress --> Complete: every row reached, none skipped - InProgress --> CompleteWithErrors: every row reached, some skipped - InProgress --> Halted: too many rows skipped, most of a small category lost,
an error it did not expect, or a shortfall the recount confirms - Halted --> InProgress: fix the cause, restart,
carry on from the cursor - Halted --> Abandoned: accept the loss, deliberately - InProgress --> Abandoned: accept the loss, deliberately - Complete --> [*] - CompleteWithErrors --> [*] - Abandoned --> [*] -``` +1. Read the reason in the refusal, the log or `--migration-status`. It names the skipped rows by reason, or the error. +2. Fix the cause. It is usually outside the migration: the body store unreachable, a certificate expired, the source or target down, or the disk full. +3. Run `--migration-retry `, then start ServiceControl. -**Four things halt a category.** +**What a retry does:** -- **Too many skips.** The skipped rows in this run pass both the percentage and the floor, which says the failures are systematic rather than incidental. -- **Most of a small category lost.** A category with fewer rows than the floor loses more than half of them. -- **An error the copy did not expect.** The error type and the cursor it stopped at are recorded. -- **A shortfall that survives the recount.** The copy came up short of the starting count, and the missing rows still exist in RavenDB but were never read. +- **It re-reads the whole category from the start**, every time, which recovers from anything. Rows already in SQL are left alone, so only the rows that failed can change. Every message body is read again, so a retry takes about as long as the first copy. +- **Historic retry operations and pending integration events** get new keys from SQL on every insert, so a re-read could not recognise rows already copied. A retry of either first deletes what that category copied, in the same save as the reset. That is safe because both are required, so ServiceControl has not opened and the target started empty. +- **The event log cannot be retried.** It also gets new keys, but it copies after ServiceControl opens, when copied rows cannot be told apart from live ones. It can only be abandoned. -A host being shut down is none of these: it leaves the category in progress, to be picked up from the cursor next time. Nor is a second host writing to the same checkpoint, which is refused so that the other host's progress stands. +**To abandon:** -**A halt stops that category and nothing else.** The remaining categories still run, with one exception: a category that must follow the halted one goes to blocked rather than running early, which is how group comments stay behind the unresolved failed messages their groups are built from. A blocked category is not a failure and needs no separate action, since clearing the halt clears the block on the next restart. +- **A required category can be abandoned once it is Failed or has started copying.** One that never ran cannot, so nothing required is given up blind. If RavenDB is gone before it starts, point back or start over: ServiceControl has not opened, so nothing being served is lost. +- **An optional category can be abandoned at any time.** That is how to stop a big optional copy short, or get out when RavenDB is gone. +- **Abandoning cascades to categories whose rows hang off it**, and the command names each: throughput history (it points at licensing endpoint records) and endpoint settings (it belongs to known endpoints). Group comments only waits for unresolved failed messages, so it is not abandoned with them. +- **Some skips no retry can fix**: a non-GUID id, a failed message with no processing attempts, a key over 200 characters, a missing required value. The refusal and the dry run say which, so abandon after the first failure rather than retrying. -**What it costs depends on which category halted.** A halted optional category means the instance keeps serving traffic and that one slice of history is missing until it is resumed. A halted required category means the host stays closed, so the outage carries on until the halt is cleared or the category is abandoned. That is deliberate: opening the host is the point of no return, and it should not happen with required data left behind by accident. +> [!IMPORTANT] +> Abandon is final. Nothing goes back for an abandoned category, so it is right when the data is not worth the outage, and wrong if picked by accident. -**Clearing it:** +**Nothing happens on a restart that you did not ask for.** A container or service restart policy restarts a refused host without anyone looking, and gets the same refusal each time, without reading anything. That is loud and loses nothing. -1. Read the reason on the category, in the custom check or the status command. It names the count that tripped the threshold, or the error, and the cursor either way. -2. Fix the cause. It is usually outside the migration: the body store unreachable, a certificate expired, the source or the target down, or the disk full. -3. Restart the host with `MigrationMode=true`. The category picks up at its cursor, its run counters start again at zero, and the skips already recorded stay on the row so the totals still add up at the end. -4. Repeat only if it halts again. A restart that halts at the same point is telling you the cause is still there, and a restart that gets further has made real progress, because the rows it skipped are committed and will not be read again. +### Ingestion workers -**Or abandon it, on purpose.** Abandoning marks the category as deliberately given up rather than failed, which lets the host open and lets the migration end. It is the right answer when the data is not worth the outage, and the wrong one if it was picked by accident, because nothing goes back for an abandoned category afterwards. What it leaves behind is stated in the counts rather than guessed at. +- **An `--error-ingestion-only` worker never runs the copy.** On RavenDB it is refused, as before, because only SQL can scale out. +- **On SQL it starts only when every required category is Done or Abandoned**, whether or not its own migration setting is on. Otherwise it refuses and names each required category still outstanding. Optional categories still copying do not hold it back. +- **That check runs even on a worker someone forgot to flag**, because ingesting into a part-copied database would make going back a loss. +- **`--import-failed-errors` on SQL runs the same check**, because it writes failed messages too. +- **A worker or import that gets past the check marks the target as opened**, whether or not its own migration setting is on. Once it has written, going back to RavenDB would lose that data, so it ends the free abort just as the main host opening does. -## Dry run +> [!IMPORTANT] +> **Start workers only after the main host has opened.** One started at the same moment as the first migration start can slip in before the copy has written its first checkpoint, and nothing stops it. -Runnable before anything starts, and again later against whatever is still outstanding. It never writes to RavenDB. +### Going back to RavenDB -What it resolves and reports: +**While ServiceControl is closed**, only the copier has written to SQL and the migration has written nothing to RavenDB, which is still the source of truth. This is the *free abort*. RavenDB's own expiration still runs, though, unless you disabled it. -- Whether the source is embedded or external, and which server -- Which two RavenDB databases, and the setting each name came from -- What it found in each of them -- Rows per category, and message-body volume per category -- A duration for the window while ServiceControl is closed, as a range +- **To go back:** set `ServiceControl/Migration/Enabled=false`, point `ServiceControl/PersistenceType` back at RavenDB, and start. You lose the copy, not your data. +- **To try again later:** start from a new, empty SQL database (`--setup`). -It runs the same startup checks that gate startup, so a missing setting surfaces before a customer books an outage. +> [!IMPORTANT] +> Never reuse the old copy. It would skip every category already finished and miss anything RavenDB received since, and nothing can detect the reuse, because it looks exactly like a normal resume. +> +> **Once ServiceControl opens**, new failed messages go to SQL and there is no way back. You can only finish the migration, or abandon what is left. -If the source holds failed error imports, it says how many and advises running `--import-failed-errors` against RavenDB before the move. They are copied either way, but ones imported first arrive as ordinary failed messages rather than as imports still waiting. +**Turning migration off before every category is Done or Abandoned is refused**, not a quiet exit. That is what stops a migration ending by accident. -It counts, before anything moves, the rows the target says it would skip or merge, by reason. These include: +- **Every SQL start checks the checkpoint table, with migration on or off.** It costs nothing when nothing is outstanding, and a RavenDB instance never reaches it. +- **With nothing outstanding, nothing changes.** A SQL instance starts exactly as before, without opening RavenDB. +- **An optional category that never started and whose window is `0` is not outstanding.** It was turned off before it began. +- **With anything outstanding, the host refuses.** It names each category with its state, counts and last error, and the ways out: + - turn migration back on to finish, after `--migration-retry` for any Failed category; + - `--migration-abandon` what you are giving up on; + - point back at RavenDB, which is free only if ServiceControl never opened on SQL (the refusal says which); + - start over against a new, empty SQL database. -- Documents whose `UniqueMessageId` will not parse as a GUID -- Subscriptions that differ only in message-type version, and so merge onto one row -- Subscriptions whose message type or transport address exceeds the 200-character key limit -- Rows already past their retention period, and rows missing a value SQL requires +**Keep RavenDB until every category is Done or Abandoned.** Before ServiceControl has opened, nothing can start without it. After, ServiceControl still opens, and the optional categories wait for RavenDB to come back, or can be abandoned. -It reports no duration for the optional categories, and nothing about load on the source. +### The dry run -### When you can run the read-only commands +The dry run never writes to RavenDB, and can run again later against whatever is outstanding. It reports: -`--migration-source-report`, `--migration-verify` and `--migration-dry-run` all open the RavenDB source. **On an embedded source that means stopping the ServiceControl service first**, because a second RavenDB process cannot attach to a data directory the first one holds. Plan the dry run as part of the outage rather than as something you run the day before while the instance keeps serving traffic. On an external source, a container or RavenDB Cloud, all three run against a live instance with no interruption. +- Whether the source is embedded or external, and which server. +- Both RavenDB database names, and the setting each came from. +- Rows and body volume per category, counting only what is inside each window. +- **A range for how long ServiceControl will be closed. The range is a floor, not a promise:** it times counting the rows and sizing the bodies, not reading every row and body or writing to SQL, so the required copy can take longer. +- **A duration for each optional category**, timed from a sample of its rows and bodies plus the pause between batches. ServiceControl is open while these copy. +- The result of the same startup checks a real start runs, so a missing setting surfaces before anyone books an outage. +- How many failed error imports there are, with advice to run `--import-failed-errors` on RavenDB first. They are copied either way, but ones imported first arrive as ordinary failed messages. +- Rows that would be skipped or merged, by reason, and which of those no retry can fix. If a required category holds any, it will end Failed, so plan to abandon it after its first failure. -A containerised instance runs all three as a one-off `docker run` of the same image with the command's flag, against an external RavenDB server, as the [instructions](ravendb-to-sql-migration-instructions.md#report-on-the-source) show for the source report. It cannot use an embedded source, because the image does not ship the RavenDB server. +It says nothing about load on the source. -`--migration-status` is the exception and is deliberately so: it reads only the checkpoint table in SQL and never opens the source, so it works on every source shape at any time, including during the background copy. It is the command to use for watching progress. +**When each read-only command can run:** -## Configuration and control +| Command | What it does | Opens RavenDB? | +| --- | --- | --- | +| `--migration-source-report` | Source facts and a document count per collection | Yes | +| `--migration-dry-run` | The report above | Yes | +| `--migration-verify` | Counts RavenDB again and compares each category with the rows the copy read, with skips and merges broken out. Exits 0 only when every category is Done or Abandoned, so a script can ask whether the migration is finished. Rows the copy never read are shown, not failed | Yes | +| `--migration-status` | Each category's state, progress, last error and the commands it can take | No, so it runs any time | -- You set `MigrationMode`, and next to it whether to copy the one optional category, archived and resolved messages. A status command and a custom check report back. -- Categories are read fresh at every startup. Adding one copies it on the next restart, removing one deletes nothing. -- There is no HTTP API, no pause, no resume, no abort command, and no way to add a category to a running instance. All of those mean editing configuration and restarting. Going back during the closed window means stopping and reconfiguring, as [the one point you can go back](#the-one-point-you-can-go-back) describes. -- The checkpoint table is a record of what happened, not a control channel. -- Stopping a copy takes a restart, so it cannot be stopped in ten seconds. +- **On an embedded source, stop the ServiceControl service before any command that opens RavenDB.** A second RavenDB process cannot use a data directory the first one holds, so the dry run is part of the outage. +- **Verify reads every selected category in full**, though not the message bodies. On a big instance that takes a while, and on an embedded source that time is part of the outage too. +- **On an external source** they all run against a live instance. +- **In a container**, run each as a one-off `docker run` of the same image against an external RavenDB server. -## Out of scope +**What verify compares:** -- The audit instance, which has no EF Core persister at all, so a customer who finishes this migration is still running RavenDB for audit. This is stated up front under [Purpose](#purpose), because it changes whether the migration is worth doing at all -- The monitoring instance, which keeps its data in memory, so there is nothing to move +- **RavenDB against the rows the copy read, never against SQL as it is now.** Once ServiceControl opens, normal use changes most categories' SQL rows within minutes: ingestion, heartbeats, retries, archiving, the licensing collectors, ServiceControl's own custom checks and the retention sweep. A SQL count then measures that use, not the copy. +- **RavenDB holding more rows than the copy read means rows were never read.** Nothing writes to RavenDB once the copy starts, so it can only shrink. Verify flags these rows. They are not in SQL, so keep RavenDB. The same flag shows if something other than this migration wrote to RavenDB since, such as an old instance still running on it. +- **RavenDB holding fewer is shown, not judged.** RavenDB's expiry keeps deleting archived and resolved messages, event log items and a few unresolved messages from 6.18 or earlier, after the copy has read them. +- **The same expiry can hide rows that were never read**, because a deleted row and a missed row cancel out. A window shorter than the retention period keeps a windowed category clear of this, by the difference between the two. +- **A windowed category is counted inside the window it started with**, on RavenDB only. +- **Verify also prints how many rows SQL held for each category when it settled.** Before ServiceControl opens nothing else writes, so for a required category that is exactly what the copy left in SQL. For the two optional categories it includes what ServiceControl wrote while they copied. + +### Monitoring and settings + +| Setting | Default | What it does | +| --- | --- | --- | +| `ServiceControl/Migration/Enabled` | `false` | Turns the migration on | +| `ServiceControl/Migration/EventLogWindow` | `ServiceControl/EventRetentionPeriod` | How far back the event log copy goes. `0` turns it off | +| `ServiceControl/Migration/ArchivedAndResolvedFailedMessagesWindow` | `ServiceControl/ErrorRetentionPeriod` | How far back the archived and resolved copy goes. `0` turns it off | +| `ServiceControl/Migration/ThrottlePauseMilliseconds` | `100` | Pause between background batches | +| `ServiceControl/Migration/HaltThresholdPercent` | `5` | Percentage part of the early-stop threshold | +| `ServiceControl/Migration/HaltThresholdMinimum` | `100` | Floor part of the early-stop threshold | + +- **The source needs no new settings.** It reads the instance's existing RavenDB settings. +- **Settings are read fresh at every start**, except that an optional category's window is fixed once that category starts. +- **Progress shows in the ServicePulse activity feed, the log and `--migration-status`.** The activity feed records when the background copy starts and finishes, when a category fails, and when the source cannot be reached. A start with migration still on after every category is Done or Abandoned logs a warning to turn it off. There is no migration custom check, because custom checks are being removed from ServiceControl. +- **Decisions are two commands, run with ServiceControl stopped:** `--migration-retry` and `--migration-abandon`. There is no HTTP API, pause, resume or abort, and no way to add a category to a running instance. +- **Stopping a copy takes a restart.** + +## Glossary + +- **Category:** one kind of data copied as a unit, such as known endpoints or the event log. +- **Required category:** one ServiceControl needs the moment it opens, so it copies while ServiceControl is closed. +- **Optional category:** one copied in the background after ServiceControl opens. It can be shortened or turned off before it starts. +- **Window:** how far back an optional category copies. +- **Closed period:** the time ServiceControl is not serving traffic because the required copy is running. +- **Checkpoint:** the row in SQL recording one category's progress, saved in the same transaction as the rows it describes. +- **Cursor:** the position in RavenDB the copy has reached in a category. +- **State:** one of Copying, Done, Failed or Abandoned. +- **Copying:** not finished yet, including waiting to start, or stopped on an error that the next start will try again. +- **Done:** every row is in SQL, or was left out for a harmless reason, and the counts balance. +- **Failed:** something went wrong. Stays Failed until you run `--migration-retry` or `--migration-abandon`. +- **Abandoned:** you gave up on what a category had not copied, or on one it depends on. Final. +- **Already present:** a row the copier found in SQL already, so it left it alone. +- **Fault skip:** a row lost because something went wrong. +- **Harmless skip:** a row left behind because SQL would have removed it anyway. +- **Free abort:** going back to RavenDB at no cost, possible only while ServiceControl is still closed. +- **Dry run:** a read-only rehearsal that reports what would move, what would be skipped and how long the outage would be. + +## Further reading + +- [The instructions](ravendb-to-sql-migration-instructions.md): how an operator runs the migration. +- [How the migration is put together](ravendb-to-sql-migration-system-design.md): which class does what, and what is built so far. diff --git a/docs/migration/ravendb-to-sql-migration-system-design.md b/docs/migration/ravendb-to-sql-migration-system-design.md new file mode 100644 index 0000000000..2e8d087297 --- /dev/null +++ b/docs/migration/ravendb-to-sql-migration-system-design.md @@ -0,0 +1,289 @@ +# Migration Engine: System Architecture Design + +> [!NOTE] +> This build does not carry the whole migration yet. This page describes it as it will be when it ships. + +[The overview](ravendb-to-sql-migration-overview.md) covers behaviour and guarantees. [The instructions](ravendb-to-sql-migration-instructions.md) cover the operator procedure. This page maps both onto types, call order and assembly boundaries. + +> [!IMPORTANT] +> Startup check order and category order are load-bearing. Read the required copy and category sections before reordering checks or adding a category. + +## Overview + +```mermaid +flowchart LR + subgraph host["ServiceControl host"] + req["RequiredCopyBeforeTheHostOpens
then MigrationStartup"] + opt["OptionalCategoryCopier"] + guard["EndOfMigrationGuard"] + cmd["Migration commands"] + end + subgraph neutral["ServiceControl.Persistence"] + eng["MigrationEngine"] + end + subgraph raven["ServiceControl.Persistence.RavenDB"] + src["RavenMigrationSource
IMigrationCategoryReader per category"] + end + subgraph ef["ServiceControl.Persistence.EFCore"] + tgt["EFCoreMigrationTarget
IMigrationCategoryWriter per category"] + cp["EFMigrationCheckpointStore"] + end + req --> eng + opt --> eng + guard --> cp + eng --> src + eng --> tgt + eng --> cp + cmd --> cp + cmd --> src + cmd --> tgt + src --> rdb[("RavenDB
read only")] + tgt --> sql[("SQL Server or PostgreSQL")] + cp --> sql + tgt --> bod[("Body store")] +``` + +`MigrationEngine` is store-agnostic. It drives `IMigrationSource` and `IMigrationTarget` per category and treats the cursor as an opaque string that it hands from source to target. Progress is persisted per category as a checkpoint row in the target database, committed in the same transaction as the rows it describes. Every command except `--migration-source-report` reads the checkpoint table. + +## Assemblies + +Dependencies point inward: each persister knows only its own store, the neutral assembly knows neither, and the host is the only composition root that names both. + +| Assembly | Types | Boundary rationale | +| --- | --- | --- | +| `ServiceControl.Persistence` | `MigrationEngine`, `MigrationBodyLoader`. Contracts: `IMigrationSource`, `IMigrationTarget`, `IMigrationCheckpointStore`, `IMigrationStartupCheck`, `IMigrationTargetReadiness`, `IMigrationSourceFactory`. Data: `MigrationCategory`, `MigrationCategoryKind`, `MigrationCategoryIds`, `MigrationCategoryRegistry`, `MigrationBatch`, `MigrationRow`, `MigrationBody`, `MigrationCheckpoint`, `MigrationWriteResult`, `MigrationSkipReason`, `MigrationSourceDescription`, `MigrationSourceFact`. Rules: `MigrationCategoryStateExtensions`, `MigrationSkipReasonExtensions`, `MigrationCheckpointRules`, `HaltThreshold`. Options: `MigrationEngineOptions`, `MigrationSettings`. `MigrationCheckpointConflictException`. Read-only views: `MigrationStatusView`, `MigrationVerification`, `MigrationDryRunArithmetic` | Referenced by both persisters and the host, references neither persister, so the engine is unit-tested against in-memory fakes in `ServiceControl.UnitTests/Migration` | +| `ServiceControl.Persistence.RavenDB` | `RavenMigrationSource`, `RavenReadOnlySourceLifecycle`, `RavenDocumentStream`, `IMigrationCategoryReader`, `IWindowedMigrationCategoryReader`, one reader per category | Owns document id prefixes, sessions and the embedded server | +| `ServiceControl.Persistence.EFCore` | `EFCoreMigrationTarget`, `EFCoreMigrationTargetReadiness`, `IMigrationCategoryWriter`, `IStoreKeyedCategoryWriter`, one writer per category, `PreparedBatch`, `IMigrationSqlDialect`, `MigrationInsert`, `StoreKeyedInsert`, `EFMigrationCheckpointStore`, `MigrationCheckpointExtensions`, `EFCoreMigrationRowAssessor`, and the three target readiness checks | Owns tables, column widths and provider parameter limits | +| `ServiceControl` (host) | `RequiredCopyBeforeTheHostOpens`, `RecordHostOpenedOnTarget`, `MigrationStartup` with nested `ClosedWindowProgress`, `MigrationStartupCheckRunner`, the four host checks, `OptionalCategoryCopier`, `RequiredStartSourceOutage`, `MigrationComponent`, `EndOfMigrationGuard`, `FinishedCopyBeforeAnIngestionNodeOpens`, `MigrationProgressReporter`, the six migration commands with `MigrationOperatorCommands` and `MigrationSourceDescriptionPrinter`, `PersistenceFactory.CreateMigrationSource` and `OpenMigrationSource` | Pair support is the one question neither persister can answer, so it lives here | + +Supporting assemblies: + +- `SqlServerMigrationSqlDialect` and `PostgreSqlMigrationSqlDialect` implement `IMigrationSqlDialect` and are the only provider-specific migration code. SQL Server reads key-column collations in `Open` and inserts with `MERGE ... WITH (HOLDLOCK)`, de-duplicating the batch with `ROW_NUMBER() OVER (PARTITION BY COLLATE )`. PostgreSQL uses `INSERT ... ON CONFLICT DO NOTHING RETURNING`. Both implement `SetLicensingEndpointThroughput`. Each provider assembly carries its own `AddMigrationCheckpoints` EF migration and model snapshot, so a checkpoint schema change needs a migration in both. +- `ServiceControl.DomainEvents` holds `MigrationStarted`, `MigrationCategoryHalted`, `MigrationFinished` and `MigrationSourceUnreachable`. + +The RavenDB persister registers no `IMigrationTarget`, `IMigrationTargetReadiness` or `IMigrationCheckpointStore`, so the guards below resolve `null` on RavenDB and return. `MigrationPairIsSupportedCheck` is what rejects RavenDB as a target. + +## What runs on every SQL instance + +These run whether or not `Migration/Enabled` is set. Each one closes a path to data loss or to a migration ending unnoticed. + +- **Registrations.** `BasePersistence.RegisterDataStores`, reached through each provider's `AddPersistence`, registers `IMigrationCheckpointStore` as `EFMigrationCheckpointStore`, `IMigrationTargetReadiness` as `EFCoreMigrationTargetReadiness` and `IMigrationTarget` as `EFCoreMigrationTarget`, all singletons. +- **`EndOfMigrationGuard`.** With `Migration/Enabled` off, `RunCommand` registers a hosted service in place of `RequiredCopyBeforeTheHostOpens` that runs `MigrationStartup.RunEndOfMigrationGuard` in `StartingAsync`, before any other hosted service and before Kestrel binds. It reads every checkpoint row and refuses if any category `IsOutstanding`. `IMigrationTargetReadiness.HasHostOpened` only selects the refusal text: whether going back to RavenDB is still free. With no checkpoint store registered (RavenDB) it returns immediately. +- **`FinishedCopyBeforeAnIngestionNodeOpens`.** Registered by `ErrorIngestionOnlyCommand.BuildHost` on every ingestion-only host and by `ImportFailedErrorsCommand.BuildHost` on every SQL import, flag on or off. In `StartingAsync` it reads every checkpoint row, drops the rows the registry classifies as optional, and refuses unless every remaining row `IsFinished`. An id unknown to this build counts as required. The worker never copies: a single host owns the copy, because a second copier would race it on the checkpoint version. +- **`RecordHostOpenedOnTarget` on workers.** Every host that passes the ingestion gate also registers `RecordHostOpenedOnTarget`, flag on or off, and stamps `Migration/HostOpenedOnTarget` when checkpoint rows exist. Ingestion ends the free abort, so the end-of-migration guard must never offer a free abort after a worker has written. + +The ingestion gate sees categories the copy never reached because `MigrationStartup.RunRequiredCategories` upserts a `NotStarted` row for every category without one before copying the first. `MigrationEngine.RunCategories` does the same for the background copy. Without those rows, a copy interrupted between two categories leaves only terminal rows and the gate passes. The gate cannot see a copy that has not written its first row. + +> [!IMPORTANT] +> Workers must start only after the main host has opened. A worker started together with the first migration start can pass the gate before the first checkpoint row exists, and nothing enforces the order. + +## The required copy + +```mermaid +sequenceDiagram + participant Run as RunCommand + participant Copy as RequiredCopyBeforeTheHostOpens + participant Start as MigrationStartup + participant Checks as MigrationStartupCheckRunner + participant Source as RavenMigrationSource + participant Target as EFCoreMigrationTarget + participant Engine as MigrationEngine + + Run->>Run: registers Copy as a hosted service, then hostBuilder.Build() + Run->>Copy: StartingAsync, before any other hosted service starts + Copy->>Start: RunRequiredCopy(app.Services, settings) + Start->>Checks: MigrationPairIsSupportedCheck + Start->>Start: PersistenceFactory.CreateMigrationSource, not connected + Start->>Checks: RunChecksAndOpenTarget + Checks->>Target: Open + Start->>Start: IMigrationCheckpointStore.ReadAll + Start->>Source: Open + Start->>Checks: source.ContributedChecks(), empty for RavenDB + Start->>Engine: RunRequiredCategories under ClosedWindowProgress + Engine-->>Start: one checkpoint per category + Start->>Start: ReportWhatTheCopyLeftBehind + Start->>Start: RefuseIfAnyCategoryDidNotComplete + Start-->>Copy: returns, or throws and the host never opens + Note over Run: RecordHostOpenedOnTarget stamps the target in StartedAsync +``` + +Each position in the sequence is deliberate: + +1. **Inside host start.** `RunCommand` registers `RequiredCopyBeforeTheHostOpens` only when `Migration/Enabled` is true, before `hostBuilder.Build()`. `app.RunAsync` drives its `StartingAsync` ahead of every other hosted service, so the retention sweeper, heartbeat settings sync, throughput collectors and the API cannot run while a required category is unfinished. Running inside `RunAsync` also keeps the Windows Service Control Manager's 30-second start timeout satisfied. +2. **`MigrationPairIsSupportedCheck` first.** It needs no source, target or options. It resolves `PersistenceType` through `PersistenceManifestLibrary` and checks it against `PersistenceFactory.SqlPersistenceNames`; the source type is fixed at `PersistenceFactory.MigrationSourcePersistenceType`. RavenDB with the flag on is rejected here. +3. **`PersistenceFactory.CreateMigrationSource`** resolves the source persistence and hard-casts its configuration to `IMigrationSourceFactory`, implemented only by `RavenPersistenceConfiguration`. A non-implementing persister fails with `InvalidCastException`. No I/O happens yet. +4. **`MigrationStartup.RunChecksAndOpenTarget`**, the target half, in order: + - `OptionalCategoryWindowsAreValidCheck` runs `MigrationEngineOptions.FromSettings`, refusing a window that is not a non-negative `TimeSpan` and naming the key. A zero window removes the category from `SelectedOptionalCategoryIds`. + - `RetryHistoryDepthIsSafeCheck` is host-owned because it needs only `Settings.RetryHistoryDepth`, which keeps that value off `PersistenceSettings` and out of `/api/configuration`. + - `EFCoreMigrationTargetReadiness.ContributedChecks`: `SchemaIsCurrentCheck`, `TargetHoldsNoServiceControlDataCheck`, `BodyStorageIsWritableCheck`. The body probe is deleted after writing, so it does not skew body counts. + - `IMigrationTarget.Open`, where the SQL Server dialect loads key-column collations. + - `OptionalCategoryWindowsAreUnchangedCheck`, last because it reads checkpoint rows. It refuses a started, unfinished optional category whose configured window differs from `StartedWindowSeconds`. + + `MigrationStartupCheckRunner` short-circuits on the first refusal. +5. **`TargetHoldsNoServiceControlDataCheck`** passes immediately if a checkpoint row exists for any required category. Otherwise it refuses if any mapped table other than `MigrationCheckpoints` has a row, `Settings` included, naming tables by `GetTableName()` (`endpoint_settings`, `settings` on PostgreSQL). +6. **Source open, last.** It is the only step that can spawn a process. `RunRequiredCopy` reads the checkpoint rows and builds the engine first. `RavenReadOnlySourceLifecycle.Open` then: + - starts the embedded server, if configured; + - connects: on an external server it rejects an expired or not-yet-valid client certificate and an `https://` URL without one, then attaches `RefuseWrite` to the request pipeline before `Initialize`, so the version request is covered; + - checks an external server's version against the client; + - waits for both databases, with a 5-minute budget per database on an embedded server. The budget is only checked when RavenDB raises its own timeout, so the wait can exceed it. + + > [!IMPORTANT] + > Readers must stream by id prefix, load by id, or query a static index. `RefuseWrite` blocks writes but passes query POSTs (`/queries`, `/multi_get`, `/streams/queries`), and a dynamic query creates an `Auto/` index. Only the embedded server sets `DisableAutoIndexCreation`. +7. **Source outage tolerance.** A failed source open is tolerated only when every required category already `IsFinished`. The host logs it, `MigrationStartup.RecordSourceOutage` writes the exception (inner exception when present) as `LastError` on each selected optional category still copying, creating `NotStarted` rows where absent, the copy is skipped and the host opens. `OptionalCategoryCopier` then stays idle for that start. In every other case the host refuses and writes nothing; if a required category never started, the refusal says the only exits are pointing back at RavenDB or starting over. +8. **`source.ContributedChecks()`** runs once the source is open, because such a check needs a session. `RavenMigrationSource` returns none, and a refusal here is never tolerated. +9. **`MigrationStartup.RunRequiredCategories`** upserts `NotStarted` rows, then runs each required category under its own linked `CancellationTokenSource`, awaited with `WaitAsync` on that token. +10. **`ClosedWindowProgress`** polls the checkpoint store every 30 seconds and logs progress. When the running category has committed nothing for 30 minutes, measured from the later of `LastProgressAt` and its run start, it cancels that category's token. `WaitAsync` returns even if the call underneath ignores cancellation, the category settles `Halted` with the stall in `LastError`, and the next category runs. The abandoned call keeps running until the host refuses and the process exits. A host shutdown leaves the row `InProgress`. +11. **`CopyOrExplainWhyItStopped`** maps `MigrationCheckpointConflictException` to a message about a second instance writing the same checkpoints. +12. **`ReportWhatTheCopyLeftBehind`** logs skips per category. It offers `--migration-retry` only for an `IsFailed` row, and reports harmless skips on a `Complete` row as harmless. +13. **`RefuseIfAnyCategoryDidNotComplete`** keeps the host closed. One message lists every Failed required category with stored state, counts, skips by reason (permanent reasons flagged as not retryable through `IsPermanent`), `LastError` and both commands. A Copying required category gets a line without a command. Rollback advice comes last. +14. **`RecordHostOpenedOnTarget`** is registered beside `RequiredCopyBeforeTheHostOpens` with the flag on. In `StartedAsync` it writes `Migration/HostOpenedOnTarget` the first time a host opens on a database that holds checkpoint rows, so the marker means "a host opened since the copy began", not "a host ran here once". The marker ends the free abort. + +## The background copy + +`OptionalCategoryCopier` is a `BackgroundService` registered by `MigrationComponent` only with the flag on and only on the main instance; `--error-ingestion-only` passes its own component list. + +- Every 20 seconds it runs the optional categories in registry order, `EventLog` then `ArchivedAndResolvedFailedMessages`, skipping rows that are `IsFinished` or `IsFailed`. +- It opens its own source through `PersistenceFactory.OpenMigrationSource`, only when a category is still copying. On failure it calls `RecordSourceOutage` and idles for the rest of the start. It also idles when the required start already recorded an outage, which it learns from the in-process `RequiredStartSourceOutage` singleton (`Recorded`, `Error`) that `RunRequiredCopy` sets where it calls `RecordSourceOutage`. The checkpoint rows cannot carry this signal, because an optional row's `LastError` can survive from an earlier start. +- An exception leaves the category `InProgress` with `LastError`, untouched until the next start, so a transient failure never forces an abandon. + +Windowing: `MigrationEngineOptions.WindowFor` supplies the window and the reader applies it through `IWindowedMigrationCategoryReader`. The engine writes `StartedWindowSeconds` once, on the first `InProgress` upsert. `CopiesFrom` is `StartedAt` minus the currently configured window, so window stability across restarts depends entirely on `OptionalCategoryWindowsAreUnchangedCheck`. Verify counts from `StartedAt` minus `StartedWindowSeconds` instead, which is the same window while the category copies, because nothing stops the setting changing once it has finished. + +> [!IMPORTANT] +> Do not remove or reorder `OptionalCategoryWindowsAreUnchangedCheck`. The stored `StartedWindowSeconds` column does not protect the window by itself. + +`Migration/ThrottlePauseMilliseconds` inserts a delay between optional-category batches, skipping the first batch of each category. Required categories are never throttled. + +## Copying a batch + +`MigrationEngine.RunCategoryAsync` is the per-category loop. Every store-specific step is delegated. + +| Step | Owner | Rationale | +| --- | --- | --- | +| Load the checkpoint | `EFMigrationCheckpointStore.Read` | The table lives in the target database | +| Short-circuit terminal or Failed rows | `MigrationCategoryStateExtensions.IsFinished`, `IsFailed` | Single definition shared by every caller; the source is not read | +| Ordering | `MigrationCategory.MustFollow`, from `MigrationCategoryRegistry.All` | Data, not code. A follower whose predecessor is not `IsFinished` is upserted `Blocked` with `LastError` naming it | +| Start or resume | Engine: `NotStarted`, `Blocked`, or any row with `LastError` moves to `InProgress`, cursor kept | `Halted` never resumes by itself; only `--migration-retry` resets it | +| Batch size | `EFCoreMigrationTarget.BatchSizeFor` via the category's `IMigrationCategoryWriter` | Writers use `RowsPerStatementFor` from the dialect's parameter budget; `StoreKeyedInsert`-only writers use a fixed 500 | +| Read after cursor | `IMigrationCategoryReader`, usually via `RavenDocumentStream.ByPrefix`; single-document readers load by id; licensing settings readers use `LicensingSettingsSource` | No-tracking session. A cursor whose document no longer exists is refused, because RavenDB would otherwise resume after the missing id and skip rows. The cursor advances on every document seen | +| Fetch bodies | `MigrationBodyLoader` over `IMigrationSource.ReadBody` | Rows with a null body only, 8 concurrent, `MigrationEngine.MaxBodyReadAttempts` (3) with `MigrationEngineOptions.BodyRetryBackoff` (200 ms). `IsDefect` exceptions such as `NotSupportedException` are not retried and reach the batch catch. A missing attachment is skipped as `BodyUnreadable` after one attempt, since ingestion writes an attachment for every failed message. Body skips stay off the checkpoint until the write commits | +| Map documents to rows | `IMigrationCategoryWriter.Prepare`, returning `PreparedBatch` | Only the writer knows `NOT NULL` columns, key caps and what the product deletes anyway. Rows are inserted later inside the target transaction. Body-carrying writers write external bodies here, before the transaction, only for messages the target does not hold, so a crash leaves orphan bodies, never dangling references | +| Insert and checkpoint | `EFCoreMigrationTarget.Write`: `AccountForEveryRow`, then in one transaction under the execution strategy `PreparedBatch.Insert`, `AlreadyPresentIn`, `MigrationCheckpoint.Extend`, `MigrationCheckpointExtensions.UpsertCheckpoint` | Most writers use the dialect's `InsertMissing` built on `MigrationInsert`. Store-keyed rows use `StoreKeyedInsert`, because `MigrationInsert` rejects store-generated keys. Throughput days use `SetLicensingEndpointThroughput`. Each attempt clears the change tracker. Skips and merges are logged after commit, and `Extend` adds the batch's `PreparedBatch.Merges` to `MergedCount`. See [progress and checkpoints](ravendb-to-sql-migration-overview.md#progress-and-checkpoints) | +| Post-commit reconciliation | Engine compares `Write`'s saved checkpoint with the counts it restated | Mismatch settles `Halted`, either kind | +| Early stop | `HaltThreshold.Exceeded` on per-run counters | Fault skips (reasons not `IsBenign`, plus body skips) must exceed both `Migration/HaltThresholdPercent` and `Migration/HaltThresholdMinimum`. Counters are not persisted, so the threshold is per run | +| Settle | `MigrationEngine.Settle`, after `IMigrationTarget.Count` at the end of the source | Saves the category's SQL row count on the row it settles, as `TargetCountAtSettle`. A halt saves without it, because the target may be what failed, and `--migration-abandon` reads it later. Logs before saving, because the checkpoint store shares the target database and a failed save would hide the cause | + +Body placement lives in `FailedMessagesWriter`, shared by `ArchivedAndResolvedFailedMessagesWriter`. `MessageBodyClassifier.Classify` decides placement and `FailedMessageRowMapper.SetBody` fills the columns: text up to `MaxBodySizeToStore` (100 KB default) inline; longer text inline prefix plus external copy; binary, invalid UTF-8 or NUL-containing text external only; empty bodies not stored. External writes go through `IBodyStoragePersistence.WriteBody`, which owns compression and the filesystem, Azure Blob or S3 backend, and complete before the transaction opens. + +Invariants enforced across class boundaries: + +- `MigrationCheckpoint.Extend` throws unless the per-`MigrationSkipReason` counts sum exactly to the skipped count. `EFCoreMigrationTarget.AccountForEveryRow` throws unless prepared rows plus skips equal the batch size, catching rows dropped during preparation. `AlreadyPresentIn` throws when a writer over-reports. Every row is copied, skipped or already present. +- `MigrationSkipReasonExtensions.IsBenign` is fixed: `PastRetention`, `BlankGroupComment`. Everything else is a fault, `Unknown` included. +- `PreparedBatch.Merges` reports rows whose key collides with another under the target's collation. `EndpointSettingsWriter` detects them with `IMigrationSqlDialect.KeyComparer` against the batch and the existing table, and `EFCoreMigrationTarget` logs one warning per merge naming both keys and the one kept. The dropped row counts as already present, and also in `MergedCount`, the part of already present that merges explain. Every writer whose key can fold two source rows into one reports them here, or `MergedCount` understates. +- `TargetCountAtSettle` is read with `IMigrationTarget.Count` just before the save that settles a category at the end of the source, or that abandons it, and is never updated afterwards. For a required category it is exactly what the copy left in SQL: the host is closed, `TargetHoldsNoServiceControlDataCheck` started the target empty, and a second copier would move the row's version so the settle save is refused. That is why the count needs no shared transaction with the save. For an optional category it also includes what ServiceControl wrote while the copy ran. + +## Category states + +`MigrationCategoryState` has seven stored values, surfaced as four operator states: + +| Operator state | Stored | Meaning | +| --- | --- | --- | +| Copying | `NotStarted`, `InProgress`, `Blocked` | Not terminal, including a follower waiting on its predecessor and an optional category stopped by an error the next start retries | +| Done | `Complete` | Every row copied or benignly skipped, counts reconciled | +| Failed | `Halted`, `CompleteWithErrors` | Stays Failed until `--migration-retry` or `--migration-abandon` | +| Abandoned | `Abandoned` | Terminal, set by the operator directly or by cascade | + +The required-copy refusal prints both forms for a Failed row (`EndpointSettings is Failed (Halted)`) and the stored state alone for a Copying row. Retry and abandon refusals print the operator state only. + +Early stops during a run: + +- Threshold exceeded or post-commit mismatch: `Halted`, either kind. +- Required category exception or 30-minute stall: `Halted`. +- Optional category exception: stays `InProgress` with `LastError`. + +Settlement at end of source, in order: + +1. Rows read this run differ from rows processed: `Halted`. +2. Any fault skip in the checkpoint: `CompleteWithErrors`. +3. Otherwise: `Complete`. + +`MigrationCategoryStateExtensions.IsFinished` (`Complete` or `Abandoned`) and `IsFailed` (`Halted` or `CompleteWithErrors`) are the single definitions, used by the engine, the required-copy refusal, the background copier, the end-of-migration guard, the ingestion gate, status, verify and both operator commands. The host opens when every required category `IsFinished`; the migration is finished when every category is. + +`MigrationCheckpointRules` holds the rules shared by the guard, the commands and status: + +- `IsOutstanding`: not finished, except an optional `NotStarted` row whose window is `0`. +- `CanBeRetried`: an `IsFailed` row, or an optional `NotStarted` or `InProgress` row with `LastError` (which covers `RecordSourceOutage` rows). Never `Blocked` or a missing row. `MigrationOperatorCommands.Retry` rejects `EventLog` before consulting it. +- `CanBeAbandoned`: a required row that `IsFailed` or is `InProgress` with copied plus skipped plus already-present above 0, treating an id unknown to the registry as required; an optional category in any state except `Complete` or `Abandoned`, including no row. + +Checkpoint schema: one row per category in `MigrationCheckpoints` (`migration_checkpoints` on PostgreSQL), primary key `CategoryId`. Columns: `State`, `Cursor`, `CopiedCount`, `SkippedCount`, `AlreadyPresentCount`, `MergedCount`, `SkipReasons` (JSON keyed by enum name; unknown names deserialize to `Unknown`, a fault), `StartedAt`, `LastProgressAt`, `SettledAt`, `LastError`, `StartedWindowSeconds`, `TargetCountAtSettle` (null until the category first settles) and `Version`, the optimistic concurrency token. No attempt counter and no decision record. `AddMigrationCheckpoints` creates the table in both providers and has not been released, so `TargetCountAtSettle` replaces the never-written `SourceTotal` column inside it, and `MergedCount` is added there too, rather than in a second migration. Run-level facts are `Settings` rows under `Migration/`, such as `Migration/HostOpenedOnTarget`. + +## Commands + +Every migration command except the source report builds a host with `AddPersistence` and never calls `StartAsync`, so no hosted service runs, including the retention sweeper and the integration event dispatcher. The source report builds no host and calls `PersistenceFactory.OpenMigrationSource` directly. + +| Command | Class | Opens RavenDB | Reads and writes | +| --- | --- | --- | --- | +| `--migration-source-report` | `MigrationSourceReportCommand` | Yes | `IMigrationSource.Describe` and `Inventory` through `MigrationSourceDescriptionPrinter`; the copy never calls either. Each `MigrationSourceFact` carries its originating setting key. `Inventory` covers every collection, including uncopied ones | +| `--migration-dry-run` | `MigrationDryRunCommand`, `MigrationDryRunArithmetic` | Yes | `MigrationPairIsSupportedCheck`, then `MigrationStartup.RunChecksAndOpen` (both halves), so it refuses wherever a real start would. Reads `Count`, `Read`, `ReadBody`, `ReadBodyVolume`, `Inventory`, `Describe`, `IMigrationTarget.Assess` (via `EFCoreMigrationRowAssessor`) and `BatchSizeFor`. Writes no row; the body probe is the only write. Exit code 0 even when the report shows problems; a refused check throws | +| `--migration-status` | `MigrationStatusCommand`, `MigrationStatusView` | No | `IMigrationCheckpointStore.ReadAll` and `MigrationEngineOptions.FromSettings`. Read-only, safe while ServiceControl runs | +| `--migration-verify` | `MigrationVerifyCommand`, `MigrationVerification` | Yes | `ReadAll`, `FromSettings`, then `IMigrationSource.Count` per selected category, judged against the rows the copy read (below). A windowed category is counted from `CopiesFrom` set to `StartedAt` minus `StartedWindowSeconds`, the window the copy used, on the source side only. Never calls `IMigrationTarget.Count`. Streams every selected category, though not the bodies, so on an embedded source ServiceControl is stopped first. Exit code 1 only when a category `IsOutstanding`; rows never read on a Done category are reported, not failed | +| `--migration-retry ` | `MigrationRetryCommand`, `MigrationOperatorCommands` | No | `Retry` checks `CanBeRetried` and builds the reset row; `IMigrationTarget.ResetForRetry` persists it | +| `--migration-abandon ` | `MigrationAbandonCommand`, `MigrationOperatorCommands` | No | Checks `CanBeAbandoned`, then writes `Abandoned` rows through the checkpoint store only | + +`MigrationOperatorCommands.Retry` builds the reset row without I/O: state `NotStarted`; `Cursor`, the four counts, `SkipReasons`, `SettledAt`, `TargetCountAtSettle` and `LastError` cleared; `StartedAt` and `StartedWindowSeconds` kept. Change retry semantics there. `ResetForRetry` only persists the row it receives, in one transaction under the execution strategy. Store-keyed categories cannot be re-read idempotently, so their writers implement `IStoreKeyedCategoryWriter.DeleteCopiedRows` and the reset deletes their rows in the same transaction: + +- `RetryOperationsWriter`: `HistoricRetryOperations`, then `UnacknowledgedRetryOperations`. +- `PendingIntegrationEventsWriter`: `ExternalIntegrationDispatchRequests`. + +Both are required categories, so the host has not opened and the target started empty. `EventLog` is also store-keyed but copies after the host opens, when its rows are indistinguishable from live ones, so it is abandon-only. + +`--migration-abandon` first re-saves the named row unchanged, when it exists, to bump its version and fence off a concurrent copy. It then abandons dependants along `MigrationCategory.MustFollowIsDataLink` (`EndpointSettings` after `KnownEndpoints`, `LicensingThroughput` after `LicensingEndpoints`), deepest first, skipping `Complete` and `Abandoned` dependants and inserting `Abandoned` rows for dependants without one, and saves the named row last. `GroupComments` follows the unresolved failed messages for ordering only and is not cascaded. An optional category with no row gets a new `Abandoned` row. Every row it saves `Abandoned` carries `TargetCountAtSettle`, read just before the save, as a settle does. No domain event is raised, because ServiceControl is stopped. + +Both operator commands rely on `MigrationCheckpointConflictException` to detect a race with a running copy. + +`MigrationVerification` judges one thing per category: RavenDB now against the rows the copy read, `CopiedCount` plus `SkippedCount` plus `AlreadyPresentCount`. That sum is exactly the rows the last complete pass read, body skips included: `RunCategoryAsync` settles `Halted` when the rows a run read differ from the rows the target accounted for, a crash resume adds to the same totals, and a retry clears them. Once the source opens RavenDB only shrinks, because `RefuseWrite` refuses every client write and only RavenDB's expiry deletes. + +- **RavenDB higher than the sum** is rows the copy never read. Verify flags them, by category. +- **RavenDB lower than the sum** is rows RavenDB removed after the copy read them, normal for the event log, the archive and a few unresolved messages with a stale expiry. Printed, not judged. +- **Expiry can hide the same number of unread rows** in the categories it touches, so a flagged number is a floor. +- **The count streams the same reader the copy used** (`RavenMigrationSource.Count`), so verify finds rows one run missed, such as a resume that skipped a range or a stream that ended early. It cannot find a reader that always misses the same rows, or a source pointed at the wrong database. +- **No live SQL count is compared with anything.** Once the host opens, normal use rewrites most categories' tables within minutes. `MergedCount` and `TargetCountAtSettle` are printed beside the sum instead. +- **An abandoned category's sum covers its last pass only.** Rows an earlier pass copied before a retry are in `TargetCountAtSettle` but not in the sum, so RavenDB minus the sum overstates what was never copied. + +**Rows never read on a Done category are accepted, not fixed.** Verify prints `ROWS NOT READ`, names the category under the table and tells the operator to keep RavenDB, because those rows exist only there. The exit code ignores them, neither operator command takes a Done row, and nothing recounts RavenDB before the host opens: the misses verify can see are particular to one run and rare, and keeping RavenDB is the remedy. + +## Readers and writers + +Readers live at `ServiceControl.Persistence.RavenDB/DataMigration/Readers/Reader.cs` and are keyed in `RavenMigrationSource`'s reader dictionary. Writers live at `ServiceControl.Persistence.EFCore/DataMigration/Writers/Writer.cs` and are keyed in the `FrozenDictionary` initialised in `EFCoreMigrationTarget`. `MigrationCategoryCoverageTests` asserts that every registry category has both. + +Every reader derives from `MigrationCategoryReader` and every writer from `MigrationCategoryWriter`, and the two for one category take the same `TDocument`, the type of each row's `Document`. The reader base streams whole documents from the primary database through `WholeDocuments`, and exposes `Lifecycle` to a reader that projects or reads the throughput database itself. The writer base casts every row once and hands `PrepareDocuments` the typed documents, so a row of another type fails the batch with an `InvalidCastException` naming the category, the row and both types. Four tests hold this: `MigrationCategoryCoverageTests.Every_reader_and_its_writer_agree_on_the_document_type`, `MigrationCategoryReadersTests.Every_reader_derives_from_the_typed_base`, `MigrationCategoryWritersTests.Every_writer_derives_from_the_typed_base` and `MigrationCategoryWritersTests.Every_writer_names_its_category_and_the_row_when_a_document_is_not_its_type`. + +| Order | Category | Kind | Reader, writer | Notes | +| --- | --- | --- | --- | --- | +| 1 | `KnownEndpoints` | Required | `KnownEndpointsReader`, `KnownEndpointsWriter` | | +| 2 | `EndpointSettings` | Required | `EndpointSettingsReader`, `EndpointSettingsWriter` | Data link to `KnownEndpoints`. Copies every setting, including one for an unknown endpoint, which `HeartbeatEndpointSettingsSyncHostedService` removes once the host opens. Reports case merges on a case-insensitive name column | +| 3 | `MessageRedirects` | Required | `MessageRedirectsReader`, `MessageRedirectsWriter` | | +| 4 | `Subscriptions` | Required | `SubscriptionsReader`, `SubscriptionsWriter` | Rows differing only in message-type version merge | +| 5 | `NotificationSettings` | Required | `NotificationSettingsReader`, `NotificationSettingsWriter` | One `Settings` row via `SettingRowMapper` | +| 6 | `TrialEndDate` | Required | `TrialEndDateReader`, `TrialEndDateWriter` | One `Settings` row via `SettingRowMapper` | +| 7 | `RetryOperations` | Required | `RetryOperationsReader`, `RetryOperationsWriter` | Two tables: historic rows store-keyed via `StoreKeyedInsert`, unacknowledged rows via `InsertMissing`, which sizes the batch. Deletes its rows on retry | +| 8 | `LicensingEndpoints` | Required | `LicensingEndpointsReader`, `LicensingEndpointsWriter` | Throughput database | +| 9 | `LicensingThroughput` | Required | `LicensingThroughputReader`, `LicensingThroughputWriter` | Throughput database. Data link to `LicensingEndpoints`. `SetLicensingEndpointThroughput` sets each day's count rather than incrementing, so re-reads are idempotent | +| 10 | `LicensingReportMasks` | Required | `LicensingReportMasksReader`, `LicensingReportMasksWriter` | Throughput database via `LicensingSettingsSource`. One `Settings` row via `SettingRowMapper` | +| 11 | `LicensedEndpointDetails` | Required | `LicensedEndpointDetailsReader`, `LicensedEndpointDetailsWriter` | Throughput database via `LicensingSettingsSource`. One `Settings` row via `SettingRowMapper` | +| 12 | `UnresolvedAndRetryIssuedFailedMessages` | Required | `FailedMessagesReader`, `FailedMessagesWriter` | Carries bodies. `FailedMessageStream` projection, `FailedMessageRowMapper` | +| 13 | `PendingIntegrationEvents` | Required | `PendingIntegrationEventsReader`, `PendingIntegrationEventsWriter` | Store-keyed via `StoreKeyedInsert`. Deletes its rows on retry | +| 14 | `CustomChecks` | Required | `CustomChecksReader`, `CustomChecksWriter` | | +| 15 | `FailedErrorImports` | Required | `FailedErrorImportsReader`, `FailedErrorImportsWriter` | Carries bodies inline in `Read`, never calls `ReadBody` | +| 16 | `GroupComments` | Required | `GroupCommentsReader`, `GroupCommentsWriter` | Ordered after `UnresolvedAndRetryIssuedFailedMessages`, not a data link. Skips blanks as `BlankGroupComment` | +| 1 | `EventLog` | Optional | `EventLogReader`, `EventLogWriter` | Windowed, store-keyed, abandon-only. Skips as `PastRetention` | +| 2 | `ArchivedAndResolvedFailedMessages` | Optional | `ArchivedAndResolvedFailedMessagesReader`, `ArchivedAndResolvedFailedMessagesWriter` | Windowed, carries bodies, shares `FailedMessagesWriter` preparation. Skips as `PastRetention` | + +## Progress reporting + +There is no migration custom check: custom checks are being removed from ServiceControl, so progress uses the activity feed, the log and `--migration-status` only. + +- `MigrationProgressReporter` raises `MigrationStarted`, `MigrationCategoryHalted` (for both Failed states), `MigrationFinished` (carrying `CategoriesAbandoned`) and `MigrationSourceUnreachable`, mapped into the ServicePulse activity feed by their `EventLogMappingDefinition` classes. It logs and swallows publish failures but propagates cancellation. All four are raised from `OptionalCategoryCopier`, so they cover the background copy only, and a Failed required category produces no activity-feed entry; its record is the startup refusal and the log. +- `MigrationSourceUnreachable` is raised once per start when `OptionalCategoryCopier` idles on a source outage, its own failed open or one the required start recorded, carrying the error and the optional categories left waiting. +- On its first pass, `OptionalCategoryCopier` logs a warning to set `Migration/Enabled` to `false` and restart when `MigrationCheckpointRules.FinishedButStillEnabled` holds: no registry category `IsOutstanding`. `--migration-status` prints the same line. `IsOutstanding` rather than `IsFinished`, so an optional category turned off with a `0` window does not hold the warning back. +- `ClosedWindowProgress` logs running required categories every 30 seconds while the host is closed. +- `--migration-status` reports state, progress, last error and available commands per category at any time, and whether migration can be turned off. diff --git a/src/ServiceControl.AcceptanceTests.PostgreSql/ServiceControl.AcceptanceTests.PostgreSql.csproj b/src/ServiceControl.AcceptanceTests.PostgreSql/ServiceControl.AcceptanceTests.PostgreSql.csproj index 718b22bc78..a55fcba55b 100644 --- a/src/ServiceControl.AcceptanceTests.PostgreSql/ServiceControl.AcceptanceTests.PostgreSql.csproj +++ b/src/ServiceControl.AcceptanceTests.PostgreSql/ServiceControl.AcceptanceTests.PostgreSql.csproj @@ -12,6 +12,9 @@ + + + @@ -20,6 +23,7 @@ + @@ -31,6 +35,9 @@ + + + diff --git a/src/ServiceControl.AcceptanceTests.SqlServer/ServiceControl.AcceptanceTests.SqlServer.csproj b/src/ServiceControl.AcceptanceTests.SqlServer/ServiceControl.AcceptanceTests.SqlServer.csproj index cbc66d5b4a..0d67353174 100644 --- a/src/ServiceControl.AcceptanceTests.SqlServer/ServiceControl.AcceptanceTests.SqlServer.csproj +++ b/src/ServiceControl.AcceptanceTests.SqlServer/ServiceControl.AcceptanceTests.SqlServer.csproj @@ -12,6 +12,9 @@ + + + @@ -20,6 +23,7 @@ + @@ -31,6 +35,9 @@ + + + diff --git a/src/ServiceControl.AcceptanceTests/Recoverability/When_hosting_error_ingestion_only.cs b/src/ServiceControl.AcceptanceTests/Recoverability/When_hosting_error_ingestion_only.cs index 092652e7d9..3285728bd5 100644 --- a/src/ServiceControl.AcceptanceTests/Recoverability/When_hosting_error_ingestion_only.cs +++ b/src/ServiceControl.AcceptanceTests/Recoverability/When_hosting_error_ingestion_only.cs @@ -29,10 +29,9 @@ namespace ServiceControl.AcceptanceTests.Recoverability using ServiceControl.Infrastructure; using ServiceControl.MessageFailures; using ServiceControl.Operations; + using ServiceControl.Persistence.DataMigration; using ServiceControl.Persistence.EFCore.DbContexts; using ServiceControl.Persistence.EFCore.Entities; - using ServiceControl.Persistence.EFCore.Infrastructure; - using ServiceControl.Recoverability; using ServiceControl.Transports; class When_hosting_error_ingestion_only : AcceptanceTest @@ -97,9 +96,36 @@ public async Task Should_default_the_concurrency_to_10_when_none_is_configured() "InternalCustomChecksHostedService", // reports this node's ingestion health to the database "MetricsReporterHostedService", "HealthCheckPublisherHostedService", // inert, no IHealthCheckPublisher is registered - "ExternalIntegrationRequestsDataStore" // its drain is inert here, nothing calls Subscribe + "ExternalIntegrationRequestsDataStore", // its drain is inert here, nothing calls Subscribe + "FinishedCopyBeforeAnIngestionNodeOpens" // keeps this node out of a database a copy has not finished filling ]; + [Test] + public async Task Should_refuse_to_start_while_a_copy_into_the_database_is_unfinished() + { + var settings = await CreateSettings(); + + await new SetupCommand().Execute(new HostArguments([]), settings); + + var host = ErrorIngestionOnlyCommand.BuildHost(settings); + + try + { + await host.Services.GetRequiredService().Upsert( + new MigrationCheckpoint(MigrationCategoryIds.KnownEndpoints, MigrationCategoryState.Halted, null, 0, 0, null, null, null, null, null, "The source stopped responding.")); + + var exception = Assert.ThrowsAsync(() => host.StartAsync()); + + Assert.That(exception.Message, Does.Contain("has not finished") + .And.Contain(MigrationCategoryIds.KnownEndpoints) + .And.Contain("--migration-abandon KnownEndpoints")); + } + finally + { + await host.DisposeAsync(); + } + } + [Test] public void Should_refuse_to_start_against_unsupported_storage() { @@ -204,6 +230,7 @@ public async Task Should_serve_health_over_https_with_the_configured_certificate await new SetupCommand().Execute(new HostArguments([]), settings); host = ErrorIngestionOnlyCommand.BuildHost(settings); + host.Urls.Clear(); host.Urls.Add("https://127.0.0.1:0"); await host.StartAsync(); diff --git a/src/ServiceControl.Migration.AcceptanceTests/.editorconfig b/src/ServiceControl.Migration.AcceptanceTests/.editorconfig new file mode 100644 index 0000000000..da44eb13fb --- /dev/null +++ b/src/ServiceControl.Migration.AcceptanceTests/.editorconfig @@ -0,0 +1,6 @@ +[*.cs] + +# Justification: Test project +dotnet_diagnostic.CA2007.severity = none +dotnet_diagnostic.PS0013.severity = none +dotnet_diagnostic.PS0018.severity = none diff --git a/src/ServiceControl.Migration.AcceptanceTests/MigrationAcceptanceTest.cs b/src/ServiceControl.Migration.AcceptanceTests/MigrationAcceptanceTest.cs new file mode 100644 index 0000000000..dd194fd221 --- /dev/null +++ b/src/ServiceControl.Migration.AcceptanceTests/MigrationAcceptanceTest.cs @@ -0,0 +1,448 @@ +namespace ServiceControl.Migration.AcceptanceTests; + +using System; +using System.Collections.Generic; +using System.Globalization; +using System.Linq; +using System.Net.Http; +using System.Net.Http.Json; +using System.Threading; +using System.Threading.Tasks; +using Microsoft.AspNetCore.Builder; +using Microsoft.EntityFrameworkCore; +using Microsoft.Extensions.DependencyInjection; +using Microsoft.Extensions.Logging.Abstractions; +using NServiceBus.Extensibility; +using NServiceBus.Transport; +using NUnit.Framework; +using ServiceControl.Persistence.EFCore; +using Particular.ServiceControl.Hosting; +using ServiceBus.Management.Infrastructure.Settings; +using ServiceControl.AcceptanceTesting.InfrastructureConfig; +using ServiceControl.Hosting.Commands; +using ServiceControl.Infrastructure.WebApi; +using ServiceControl.Migration.Checks; +using ServiceControl.Operations; +using ServiceControl.Persistence; +using ServiceControl.Persistence.DataMigration; +using ServiceControl.Persistence.EFCore.Abstractions; +using ServiceControl.Persistence.EFCore.DbContexts; +using ServiceControl.Persistence.EFCore.Entities; +using ServiceControl.Persistence.EFCore.Infrastructure; +using ServiceControl.Persistence.Infrastructure; +using ServiceControl.Persistence.Tests; +using ServiceControl.Persistence.UnitOfWork; +using TestHelper; +// Aliased rather than imported: Raven.Client.Documents carries LINQ extensions that collide with EF Core's. +using IDocumentStore = Raven.Client.Documents.IDocumentStore; + +// No base class: the persistence test bases sit in projects that cannot see RavenDB. +abstract class MigrationAcceptanceTest +{ + protected const string EndpointSettingsUrl = "api/endpointssettings"; + + readonly AcceptanceTestStorageConfiguration StorageConfiguration = new(); + readonly List setVariables = []; + CancellationTokenSource hostCancellation; + Task runningHost; + + ServiceProvider targetServices; + IMigrationCheckpointStore targetCheckpointStore; + + protected Settings Settings { get; private set; } + protected HttpClient HttpClient { get; private set; } + protected (string ServerUrl, string PrimaryDatabase, string ThroughputDatabase) Source { get; private set; } + protected IDocumentStore SourceStore { get; private set; } + + protected IMigrationTarget Target { get; private set; } + + protected InMemoryBodyStoragePersistence RecordedBodies { get; } = new(); + + [SetUp] + public async Task SetUp() + { + Source = await MigrationSourceServer.CreateDatabases(); + SourceStore = await (await MigrationSourceServer.GetInstance()).Connect(); + + SetSourceVariable("SERVICECONTROL_RAVENDB_CONNECTIONSTRING", Source.ServerUrl); + SetSourceVariable("SERVICECONTROL_RAVENDB_DATABASENAME", Source.PrimaryDatabase); + SetSourceVariable("LICENSINGCOMPONENT_RAVENDB_THROUGHPUTDATABASENAME", Source.ThroughputDatabase); + SetSourceVariable("SERVICECONTROL_ERRORRETENTIONPERIOD", "10.00:00:00"); + SetSourceVariable("SERVICECONTROL_MIGRATION_ENABLED", "true"); + + // Settings.Port has no public setter, so the port has to come in as the environment variable a customer would set. + SetSourceVariable("SERVICECONTROL_PORT", PortUtility.GetAssignedOrAvailablePort(33500).ToString(CultureInfo.InvariantCulture)); + + var transport = new ConfigureEndpointLearningTransport(); + + Settings = new Settings( + transportType: transport.TypeName, + persisterType: StorageConfiguration.PersistenceType, + forwardErrorMessages: false, + errorRetentionPeriod: TimeSpan.FromDays(10)) + { + TransportConnectionString = transport.ConnectionString + }; + + await StorageConfiguration.CustomizeSettings(Settings); + await new SetupCommand().Execute(new HostArguments([]), Settings); + + HttpClient = new HttpClient { BaseAddress = new Uri(Settings.RootUrl) }; + + var targetServiceCollection = new ServiceCollection(); + targetServiceCollection.AddLogging(); + targetServiceCollection.AddPersistence(Settings); + targetServiceCollection.AddSingleton(RecordedBodies); + targetServices = targetServiceCollection.BuildServiceProvider(); + + Target = targetServices.GetRequiredService(); + targetCheckpointStore = targetServices.GetRequiredService(); + await Target.Open(); + } + + // The variables go first because they are process wide: a failure in the cleanup below must not leave them set for the next test. + [TearDown] + public async Task TearDown() + { + foreach (var name in setVariables) + { + Environment.SetEnvironmentVariable(name, null); + } + + try + { + await StopHost(); + } + finally + { + HttpClient?.Dispose(); + SourceStore?.Dispose(); + + if (targetServices is not null) + { + await targetServices.DisposeAsync(); + } + + await StorageConfiguration.Cleanup(); + } + } + + protected void SetSourceVariable(string name, string value) + { + Environment.SetEnvironmentVariable(name, value); + setVariables.Add(name); + } + + protected async Task SeedSource(string database, params (string Id, object Document)[] documents) + { + using var session = SourceStore.OpenAsyncSession(database); + + foreach (var (id, document) in documents) + { + await session.StoreAsync(document, id); + } + + await session.SaveChangesAsync(); + } + + protected Task SeedSource(params (string Id, object Document)[] documents) => + SeedSource(Source.PrimaryDatabase, documents); + + protected Task SeedSourceEndpointSettings(params (string Name, bool TrackInstances)[] settings) => + SeedSource([.. settings.Select(setting => ($"EndpointSettings/{DeterministicGuid.MakeId(setting.Name)}", (object)new EndpointSettings { Name = setting.Name, TrackInstances = setting.TrackInstances }))]); + + // The target copies an endpoint setting only when that endpoint is known, so seed the endpoint too. + protected Task SeedSourceKnownEndpoints(params string[] names) => + SeedSource([.. names + .Select(name => new KnownEndpoint { EndpointDetails = new EndpointDetails { Name = name, HostId = Guid.NewGuid(), Host = "HOST01" } }) + .Select(endpoint => ($"KnownEndpoints/{endpoint.EndpointDetails.GetDeterministicId()}", (object)endpoint))]); + + protected Task OpenSource(CancellationToken cancellationToken = default) => + PersistenceFactory.OpenMigrationSource(Settings, cancellationToken); + + static IReadOnlyCollection sourceSupportedCategoryIds; + + // SupportedCategoryIds answers before Open and cannot change across it, so one source built once for the + // whole run answers it for every test instead of loading the persister assembly again per test. + protected async Task> SourceSupportedCategoryIds() + { + if (sourceSupportedCategoryIds is null) + { + await using var source = PersistenceFactory.CreateMigrationSource(Settings); + sourceSupportedCategoryIds = source.SupportedCategoryIds; + } + + return sourceSupportedCategoryIds; + } + + + // No build carries the whole migration until its last phase ships, so without this marker every copy is refused. + protected static Action AllowingAnUnreleasedMigration(Action customize = null) => + builder => + { + builder.Services.AddSingleton(); + customize?.Invoke(builder); + }; + + protected async Task RunHostUntilTheApiAnswers(Action customize = null) + { + await StopHost(); + + hostCancellation = new CancellationTokenSource(); + runningHost = RunCommand.Run(Settings, AllowingAnUnreleasedMigration(customize), hostCancellation.Token); + + await WaitForEndpointSettingsResponse(TimeSpan.FromMinutes(2)); + } + + protected Settings SettingsWithMigrationEnabled(bool copyEventLog = false, Action customize = null) + { + SetSourceVariable("SERVICECONTROL_MIGRATION_ENABLED", "true"); + SetSourceVariable("SERVICECONTROL_MIGRATION_EVENTLOGWINDOW", copyEventLog ? null : "0"); + SetSourceVariable("SERVICECONTROL_MIGRATION_ARCHIVEDANDRESOLVEDFAILEDMESSAGESWINDOW", "0"); + customize?.Invoke(Settings); + return Settings; + } + + protected async Task RunRequiredCopy( + Action customize = null, + bool leaveOptionalIncomplete = false, + CancellationToken cancellationToken = default) + { + SettingsWithMigrationEnabled(copyEventLog: leaveOptionalIncomplete); + + await RunHostUntilTheApiAnswers(builder => customize?.Invoke(builder.Services)); + + var copyable = MigrationStartup.CopyableCategoryIds(await SourceSupportedCategoryIds(), Target.SupportedCategoryIds); + + var copied = MigrationCategoryRegistry.All + .Where(category => category.Kind == MigrationCategoryKind.Required && copyable.Contains(category.Id)) + .Select(category => category.Id) + .ToArray(); + + await WaitUntil( + async () => (await targetCheckpointStore.ReadAll(cancellationToken)) + .Count(checkpoint => copied.Contains(checkpoint.CategoryId) && checkpoint.State.IsFinished()) == copied.Length, + "the required copy finished"); + + if (leaveOptionalIncomplete) + { + // The guarded services read IMigrationState.AnyCategoryIncomplete, which any selected category that has not finished holds true. Halted is one of those states. + await targetCheckpointStore.Upsert( + new MigrationCheckpoint("EventLog", MigrationCategoryState.Halted, null, 0, 0, null, null, null, null, null, + "Halted by the test fixture so the guarded services see an unfinished category."), + cancellationToken); + } + } + + protected async Task> WaitForEndpointSettingsResponse(TimeSpan timeout) + { + IReadOnlyList settings = null; + + await WaitUntil(async () => + { + // A host that refused to start surfaces its own exception here instead of waiting out the timeout. + if (runningHost is { IsFaulted: true }) + { + await runningHost; + } + + try + { + settings = await GetEndpointSettings(); + return true; + } + catch (HttpRequestException) + { + return false; + } + }, $"GET {EndpointSettingsUrl} answered", timeout); + + return settings; + } + + protected async Task> GetEndpointSettings() => + await HttpClient.GetFromJsonAsync>(EndpointSettingsUrl, SerializerOptions.Default); + + async Task StopHost() + { + if (runningHost is null) + { + return; + } + + await hostCancellation.CancelAsync(); + await runningHost; + hostCancellation.Dispose(); + runningHost = null; + } + + // Copied from PersistenceTestBase.WaitUntil, which this project cannot compile. + protected static async Task WaitUntil(Func> conditionChecker, string condition, TimeSpan timeout = default) + { + timeout = timeout == default ? TimeSpan.FromSeconds(10) : timeout; + + var start = DateTime.UtcNow; + + while (DateTime.UtcNow - start < timeout) + { + if (await conditionChecker()) + { + return; + } + + await Task.Delay(TimeSpan.FromMilliseconds(500)); + } + + throw new Exception($"{condition} has not been meet in defined timespan: {timeout})"); + } + + // Copied from RavenMigrationSourceTestBase.CollectBatches, which this project cannot compile. + protected static async Task> CollectBatches(IMigrationSource source, MigrationCategory category, string resumeAfter = null, int batchSize = 100) + { + var batches = new List(); + + await foreach (var batch in source.Read(category, resumeAfter, batchSize, TestContext.CurrentContext.CancellationToken)) + { + batches.Add(batch); + } + + return batches; + } + + protected IEndpointSettingsStore EndpointSettingsStore => targetServices.GetRequiredService(); + protected IMonitoringDataStore MonitoringDataStore => targetServices.GetRequiredService(); + + protected async Task CopyCategory(string categoryId, CancellationToken cancellationToken = default) + { + var category = MigrationCategoryRegistry.Find(categoryId) ?? throw new ArgumentException($"'{categoryId}' is not a migration category.", nameof(categoryId)); + var before = await targetCheckpointStore.Read(categoryId, cancellationToken); + var target = new RecordingMigrationTarget(Target); + var options = new MigrationEngineOptions(TimeSpan.Zero, MigrationSettings.DefaultHaltThresholdPercent, MigrationSettings.DefaultHaltThresholdMinimum, []) { BodyRetryBackoff = TimeSpan.Zero }; + + await using var source = await OpenSource(cancellationToken); + var engine = new MigrationEngine(source, target, targetCheckpointStore, TimeProvider.System, options, NullLogger.Instance); + + var after = await engine.RunCategoryAsync(category, cancellationToken); + + // Both reached the end of the source, so a test can assert a Failed category's skips as well as a Done one's. + if (after.State is not (MigrationCategoryState.Complete or MigrationCategoryState.CompleteWithErrors)) + { + throw new InvalidOperationException($"Copying '{categoryId}' ended {after.State}: {after.LastError}"); + } + + return new MigrationWriteResult( + after, + (int)(after.CopiedCount - (before?.CopiedCount ?? 0)), + (int)(after.SkippedCount - (before?.SkippedCount ?? 0)), + target.SkippedIds, + (int)(after.AlreadyPresentCount - (before?.AlreadyPresentCount ?? 0))); + } + + + // Makes the target look like a database a migration has touched, which is the only state the host-opened marker is stamped in. + protected async Task SeedCheckpoint(string categoryId, MigrationCategoryState state = MigrationCategoryState.Complete) + { + await using var scope = targetServices.CreateAsyncScope(); + var dbContext = scope.ServiceProvider.GetRequiredService(); + + await dbContext.UpsertCheckpoint(new MigrationCheckpoint(categoryId, state, null, 0, 0, null, null, null, null, null, null)); + } + + protected async Task DeleteCheckpoint(string categoryId) + { + await using var scope = targetServices.CreateAsyncScope(); + var dbContext = scope.ServiceProvider.GetRequiredService(); + + await dbContext.MigrationCheckpoints.Where(checkpoint => checkpoint.CategoryId == categoryId).ExecuteDeleteAsync(); + } + + protected async Task QueryTarget(Func> query) + { + using var scope = targetServices.CreateScope(); + var dbContext = scope.ServiceProvider.GetRequiredService(); + + return await query(dbContext); + } + + protected async Task GetFailedMessage(Guid uniqueMessageId) + { + var row = await QueryTarget(dbContext => dbContext.FailedMessages.AsNoTracking().SingleOrDefaultAsync(m => m.UniqueMessageId == uniqueMessageId)); + + Assert.That(row, Is.Not.Null, $"No failed message row for {uniqueMessageId}"); + + return row; + } + + protected Task FindFailedMessage(Guid uniqueMessageId) => + QueryTarget(dbContext => dbContext.FailedMessages.AsNoTracking().SingleOrDefaultAsync(m => m.UniqueMessageId == uniqueMessageId)); + + protected EFPersisterSettings EFSettings => targetServices.GetRequiredService(); + + protected Task Ingest(params IngestedFailure[] failures) => + InBatch(async unitOfWork => + { + foreach (var failure in failures) + { + await unitOfWork.Recoverability.RecordFailedProcessingAttempt(failure.Context, failure.ProcessingAttempt, failure.Groups); + } + }); + + // Ingestion takes the body and its content type from the message context, so only the context changes. + protected Task IngestWithBody(IngestedFailure failure, byte[] body, string contentType) => + InBatch(unitOfWork => + { + var headers = new Dictionary(failure.Headers) { [NServiceBus.Headers.ContentType] = contentType }; + var context = new MessageContext(failure.MessageId, headers, body, new TransportTransaction(), "receiveAddress", new ContextBag()); + + return unitOfWork.Recoverability.RecordFailedProcessingAttempt(context, failure.ProcessingAttempt, failure.Groups); + }); + + async Task InBatch(Func record) + { + await using var unitOfWork = await targetServices.GetRequiredService().StartNew(); + + await record(unitOfWork); + + await unitOfWork.Complete(TestContext.CurrentContext.CancellationToken); + } + + protected Task ReadCheckpoint(string categoryId) => + QueryTarget(async dbContext => + { + var row = await dbContext.MigrationCheckpoints.AsNoTracking().SingleOrDefaultAsync(checkpoint => checkpoint.CategoryId == categoryId); + + return row is null ? null : new MigrationCheckpoint( + row.CategoryId, row.State, row.Cursor, + row.CopiedCount, row.SkippedCount, row.SourceTotal, row.SkipReasons, + row.StartedAt, row.LastProgressAt, row.SettledAt, row.LastError, + row.AlreadyPresentCount, row.Version); + }); + + protected Task ReadSetting(string key) => + QueryTarget(dbContext => dbContext.Settings.AsNoTracking() + .Where(setting => setting.Key == key) + .Select(setting => setting.Value) + .SingleOrDefaultAsync()); + + sealed class RecordingMigrationTarget(IMigrationTarget inner) : IMigrationTarget + { + public List SkippedIds { get; } = []; + + public Task Open(CancellationToken cancellationToken = default) => inner.Open(cancellationToken); + + public Task BatchSizeFor(MigrationCategory category, CancellationToken cancellationToken = default) => inner.BatchSizeFor(category, cancellationToken); + + public async Task Write(MigrationCategory category, MigrationBatch batch, MigrationCheckpoint checkpointToExtend, CancellationToken cancellationToken = default) + { + var result = await inner.Write(category, batch, checkpointToExtend, cancellationToken); + SkippedIds.AddRange(result.SkippedIds); + return result; + } + + public Task Count(MigrationCategory category, CancellationToken cancellationToken = default) => inner.Count(category, cancellationToken); + + public IReadOnlyCollection SupportedCategoryIds => inner.SupportedCategoryIds; + + public IReadOnlyDictionary DocumentTypes => inner.DocumentTypes; + } +} diff --git a/src/ServiceControl.Migration.AcceptanceTests/MigrationSourceServer.cs b/src/ServiceControl.Migration.AcceptanceTests/MigrationSourceServer.cs new file mode 100644 index 0000000000..88a33b0003 --- /dev/null +++ b/src/ServiceControl.Migration.AcceptanceTests/MigrationSourceServer.cs @@ -0,0 +1,82 @@ +namespace ServiceControl.Migration.AcceptanceTests; + +using System; +using System.IO; +using System.Threading; +using System.Threading.Tasks; +using Microsoft.Extensions.Hosting.Internal; +using Microsoft.Extensions.Logging.Abstractions; +using NUnit.Framework; +using Raven.Client.ServerWide; +using Raven.Client.ServerWide.Operations; +using ServiceControl.RavenDB; +using TestHelper; + +// Its own copy rather than SharedEmbeddedServer: that one uses the RavenDB persister's own settings type, which this project cannot reference because it loads the persister from its manifest instead. +static class MigrationSourceServer +{ + public static async Task GetInstance(CancellationToken cancellationToken = default) + { + await startLock.WaitAsync(cancellationToken); + try + { + if (server != null) + { + return server; + } + + var dbPath = Path.Combine(TestContext.CurrentContext.WorkDirectory, "Tests", "MigrationSource"); + var logPath = Path.Combine(TestContext.CurrentContext.WorkDirectory, "Logs", "MigrationSource"); + var port = PortUtility.GetAssignedOrAvailablePort(33350); + + var configuration = new EmbeddedDatabaseConfiguration($"http://localhost:{port}", "primary", dbPath, logPath, "Operations") { RunInMemory = true }; + + server = EmbeddedDatabase.Start(configuration, lifetime); + return server; + } + finally + { + startLock.Release(); + } + } + + public static async Task Stop() + { + await startLock.WaitAsync(); + try + { + if (server is null) + { + return; + } + + await server.Stop(CancellationToken.None); + server.Dispose(); + server = null; + lifetime.StopApplication(); + } + finally + { + startLock.Release(); + } + } + + public static async Task<(string ServerUrl, string PrimaryDatabase, string ThroughputDatabase)> CreateDatabases(CancellationToken cancellationToken = default) + { + var instance = await GetInstance(cancellationToken); + var primary = $"sc_src_{Guid.NewGuid():n}"; + var throughput = $"{primary}-throughput"; + + using var store = await instance.Connect(cancellationToken); + foreach (var name in new[] { primary, throughput }) + { + await store.Maintenance.Server.SendAsync(new CreateDatabaseOperation(new DatabaseRecord(name)), cancellationToken); + } + + return (instance.ServerUrl, primary, throughput); + } + + static EmbeddedDatabase server; + static readonly ApplicationLifetime lifetime = new(new NullLogger()); + static readonly SemaphoreSlim startLock = new(1, 1); +} diff --git a/src/ServiceControl.Migration.AcceptanceTests/StopMigrationSourceServer.cs b/src/ServiceControl.Migration.AcceptanceTests/StopMigrationSourceServer.cs new file mode 100644 index 0000000000..b42d2671f2 --- /dev/null +++ b/src/ServiceControl.Migration.AcceptanceTests/StopMigrationSourceServer.cs @@ -0,0 +1,11 @@ +namespace ServiceControl.Migration.AcceptanceTests; + +using System.Threading.Tasks; +using NUnit.Framework; + +[SetUpFixture] +public class StopMigrationSourceServer +{ + [OneTimeTearDown] + public Task Teardown() => MigrationSourceServer.Stop(); +} diff --git a/src/ServiceControl.Migration.AcceptanceTests/TestMigrationTargets.cs b/src/ServiceControl.Migration.AcceptanceTests/TestMigrationTargets.cs new file mode 100644 index 0000000000..fd07ff535b --- /dev/null +++ b/src/ServiceControl.Migration.AcceptanceTests/TestMigrationTargets.cs @@ -0,0 +1,121 @@ +namespace ServiceControl.Migration.AcceptanceTests; + +using System; +using System.Collections.Generic; +using System.Linq; +using System.Threading; +using System.Threading.Tasks; +using Microsoft.AspNetCore.Builder; +using Microsoft.Extensions.DependencyInjection; +using ServiceControl.Persistence.DataMigration; + +static class TestMigrationTargets +{ + public static void ParkFirstMigrationWrite(this WebApplicationBuilder builder, TaskCompletionSource parked, Task release) => + builder.DecorateMigrationTarget(inner => new ParkingMigrationTarget(inner, parked, release)); + + public static void FailTheSecondMigrationWrite(this WebApplicationBuilder builder) => + builder.DecorateMigrationTarget(inner => new FaultingMigrationTarget(inner, failingWrite: 2)); + + public static void HaltTheEndpointSettingsCategory(this WebApplicationBuilder builder) => + builder.DecorateMigrationTarget(inner => new FaultingMigrationTarget(inner, failingWrite: 1)); + + public static void StopAfterTheFirstCategorySettles(this WebApplicationBuilder builder, CancellationTokenSource stopping) + { + var registered = builder.Services.Last(service => service.ServiceType == typeof(IMigrationCheckpointStore)); + + if (registered.ImplementationType is null) + { + throw new InvalidOperationException( + "IMigrationCheckpointStore is registered by a factory rather than by type, so the test decorator cannot rebuild the inner store. Register it as AddSingleton() or give the decorator a different seam."); + } + + builder.Services.AddSingleton(provider => + new StoppingAfterTheFirstSettleCheckpointStore((IMigrationCheckpointStore)ActivatorUtilities.CreateInstance(provider, registered.ImplementationType), stopping)); + } + + public static void DecorateMigrationTarget(this WebApplicationBuilder builder, Func decorate) + { + var registered = builder.Services.Last(service => service.ServiceType == typeof(IMigrationTarget)); + + if (registered.ImplementationType is null) + { + throw new InvalidOperationException( + "IMigrationTarget is registered by a factory rather than by type, so the test decorator cannot rebuild the inner target. Register it as AddSingleton() or give the decorator a different seam."); + } + + builder.Services.AddSingleton(provider => + decorate((IMigrationTarget)ActivatorUtilities.CreateInstance(provider, registered.ImplementationType))); + } +} + +// Throws on one EndpointSettings write, which is the only way to see what a copy does after a write has failed. +sealed class FaultingMigrationTarget(IMigrationTarget inner, int failingWrite) : IMigrationTarget +{ + int endpointSettingsWrites; + + public Task Open(CancellationToken cancellationToken = default) => inner.Open(cancellationToken); + + // Two rows a batch, so five settings take three writes and a failure on the second leaves exactly one batch committed. + public Task BatchSizeFor(MigrationCategory category, CancellationToken cancellationToken = default) => + category.Id == MigrationCategoryIds.EndpointSettings ? Task.FromResult(2) : inner.BatchSizeFor(category, cancellationToken); + + public Task Write(MigrationCategory category, MigrationBatch batch, MigrationCheckpoint checkpointToExtend, CancellationToken cancellationToken = default) => + category.Id == MigrationCategoryIds.EndpointSettings && ++endpointSettingsWrites == failingWrite + ? throw new Exception("injected mid-category failure") + : inner.Write(category, batch, checkpointToExtend, cancellationToken); + + public Task Count(MigrationCategory category, CancellationToken cancellationToken = default) => inner.Count(category, cancellationToken); + + public IReadOnlyCollection SupportedCategoryIds => inner.SupportedCategoryIds; + + public IReadOnlyDictionary DocumentTypes => inner.DocumentTypes; +} + +// Stops the copy the moment the first category settles Done or Abandoned, which is the only way to stop it between two categories. +sealed class StoppingAfterTheFirstSettleCheckpointStore(IMigrationCheckpointStore inner, CancellationTokenSource stopping) : IMigrationCheckpointStore +{ + public Task> ReadAll(CancellationToken cancellationToken = default) => inner.ReadAll(cancellationToken); + + public Task Read(string categoryId, CancellationToken cancellationToken = default) => inner.Read(categoryId, cancellationToken); + + public async Task Upsert(MigrationCheckpoint checkpoint, CancellationToken cancellationToken = default) + { + var stored = await inner.Upsert(checkpoint, cancellationToken); + + if (stored.State.IsFinished()) + { + await stopping.CancelAsync(); + throw new OperationCanceledException(stopping.Token); + } + + return stored; + } +} + +// Holds the copy open at a moment a test can observe, which is the only way to ask what the API answers mid-copy. +sealed class ParkingMigrationTarget(IMigrationTarget inner, TaskCompletionSource parked, Task release) : IMigrationTarget +{ + int writes; + + public Task Open(CancellationToken cancellationToken = default) => inner.Open(cancellationToken); + + public Task BatchSizeFor(MigrationCategory category, CancellationToken cancellationToken = default) => inner.BatchSizeFor(category, cancellationToken); + + public async Task Write(MigrationCategory category, MigrationBatch batch, MigrationCheckpoint checkpointToExtend, CancellationToken cancellationToken = default) + { + if (Interlocked.Exchange(ref writes, 1) == 0) + { + parked.SetResult(); + await release.WaitAsync(cancellationToken); + } + + return await inner.Write(category, batch, checkpointToExtend, cancellationToken); + } + + public Task Count(MigrationCategory category, CancellationToken cancellationToken = default) => inner.Count(category, cancellationToken); + + public IReadOnlyCollection SupportedCategoryIds => inner.SupportedCategoryIds; + + public IReadOnlyDictionary DocumentTypes => inner.DocumentTypes; +} diff --git a/src/ServiceControl.Migration.AcceptanceTests/When_a_copy_is_killed_mid_category.cs b/src/ServiceControl.Migration.AcceptanceTests/When_a_copy_is_killed_mid_category.cs new file mode 100644 index 0000000000..a55deaba69 --- /dev/null +++ b/src/ServiceControl.Migration.AcceptanceTests/When_a_copy_is_killed_mid_category.cs @@ -0,0 +1,182 @@ +namespace ServiceControl.Migration.AcceptanceTests; + +using System; +using System.Linq; +using System.Threading; +using System.Threading.Tasks; +using Microsoft.AspNetCore.Builder; +using Microsoft.EntityFrameworkCore; +using Microsoft.Extensions.DependencyInjection; +using Microsoft.Extensions.Hosting; +using NUnit.Framework; +using ServiceControl.Hosting.Commands; +using ServiceControl.Persistence.DataMigration; +using ServiceControl.Persistence.EFCore; + +[TestFixture] +[NonParallelizable] +class When_a_copy_is_killed_mid_category : MigrationAcceptanceTest +{ + [Test] + public void A_migrating_host_stopped_while_starting_exits_cleanly() + { + using var cancellation = new CancellationTokenSource(TimeSpan.FromMinutes(3)); + + Assert.DoesNotThrowAsync(async () => + await RunCommand.Run(Settings, AllowingAnUnreleasedMigration(StoppedWhileStarting), cancellation.Token)); + } + + [Test] + public void A_host_not_migrating_stopped_while_starting_still_fails() + { + SetSourceVariable("SERVICECONTROL_MIGRATION_ENABLED", "false"); + using var cancellation = new CancellationTokenSource(TimeSpan.FromMinutes(3)); + + Assert.CatchAsync(async () => + await RunCommand.Run(Settings, AllowingAnUnreleasedMigration(StoppedWhileStarting), cancellation.Token)); + } + + // Only a stop is a clean exit; a cancellation nobody asked for is still a failed start, migrating or not. + [Test] + public void A_migrating_host_whose_start_is_cancelled_without_a_stop_still_fails() + { + using var cancellation = new CancellationTokenSource(TimeSpan.FromMinutes(3)); + + Assert.CatchAsync(async () => + await RunCommand.Run(Settings, AllowingAnUnreleasedMigration(CancelledWithoutAStop), cancellation.Token)); + } + + [Test] + public async Task It_resumes_rather_than_starting_over() + { + await SeedSourceKnownEndpoints("A", "B", "C", "D", "E"); + await SeedSourceEndpointSettings(("A", true), ("B", true), ("C", true), ("D", true), ("E", true)); + + // Bounded rather than None: a copy that did not fail would otherwise hold this call open forever. + using var cancellation = new CancellationTokenSource(TimeSpan.FromMinutes(3)); + + var killed = Assert.ThrowsAsync(async () => + await RunCommand.Run(Settings, AllowingAnUnreleasedMigration(builder => builder.FailTheSecondMigrationWrite()), cancellation.Token)); + + Assert.That(killed, Is.Not.Null); + Assert.That(await TargetEndpointSettingsCount(), Is.EqualTo(2), "the first batch and its cursor should have committed together"); + + var refusedAgain = Assert.ThrowsAsync(async () => + await RunCommand.Run(Settings, AllowingAnUnreleasedMigration(), cancellation.Token)); + Assert.That(refusedAgain.Message, Is.EqualTo(killed.Message), "a start re-ran a Failed category the operator had not put back"); + + // The row a kill part way through the category leaves: still copying, with the first batch's cursor and counts. + var halted = await ReadCheckpoint("EndpointSettings"); + await QueryTarget(dbContext => dbContext.UpsertCheckpoint(halted with { State = MigrationCategoryState.InProgress, LastError = null, SettledAt = null })); + + await RunHostUntilTheApiAnswers(); + + var checkpoint = await ReadCheckpoint("EndpointSettings"); + + using (Assert.EnterMultipleScope()) + { + Assert.That(await GetEndpointSettingsNames(), Is.EquivalentTo(new[] { "A", "B", "C", "D", "E" }), "no gaps"); + Assert.That(checkpoint.CopiedCount, Is.EqualTo(5), "no row copied twice"); + Assert.That(checkpoint.SkippedCount, Is.Zero); + Assert.That(checkpoint.AlreadyPresentCount, Is.Zero, "a resumed run must not re-read rows the first run committed"); + } + } + + // Proves the already-present count in the test above can fail: both runs end with five correct rows, so that count is the only thing telling a resume from a fresh copy. + [Test] + public async Task Without_its_checkpoint_the_second_run_re_reads_what_the_first_already_wrote() + { + await SeedSourceKnownEndpoints("A", "B", "C", "D", "E"); + await SeedSourceEndpointSettings(("A", true), ("B", true), ("C", true), ("D", true), ("E", true)); + + using var cancellation = new CancellationTokenSource(TimeSpan.FromMinutes(3)); + + Assert.ThrowsAsync(async () => + await RunCommand.Run(Settings, AllowingAnUnreleasedMigration(builder => builder.FailTheSecondMigrationWrite()), cancellation.Token)); + + await DeleteCheckpoint("EndpointSettings"); + await RunHostUntilTheApiAnswers(); + + var checkpoint = await ReadCheckpoint("EndpointSettings"); + + Assert.That(checkpoint.AlreadyPresentCount, Is.EqualTo(2)); + } + + // A copy stopped between two categories must still list the second, or every gate reading the table finds nothing outstanding. + [Test] + public async Task A_worker_started_after_a_copy_stopped_between_categories_names_the_one_it_never_reached() + { + await SeedSourceKnownEndpoints("A", "B"); + await SeedSourceEndpointSettings(("A", true), ("B", true)); + Settings.MaximumConcurrencyLevel = 1; + + using var stopping = new CancellationTokenSource(TimeSpan.FromMinutes(3)); + Assert.CatchAsync(async () => + await RunCommand.Run(Settings, AllowingAnUnreleasedMigration(builder => builder.StopAfterTheFirstCategorySettles(stopping)), stopping.Token)); + + using var cancellation = new CancellationTokenSource(TimeSpan.FromMinutes(3)); + var worker = ErrorIngestionOnlyCommand.BuildHost(Settings, AllowingAnUnreleasedMigration()); + try + { + using (Assert.EnterMultipleScope()) + { + Assert.That((await ReadCheckpoint(MigrationCategoryIds.KnownEndpoints))?.State, Is.EqualTo(MigrationCategoryState.Complete), "the stop has to land between the two categories for this test to mean anything"); + Assert.That(async () => await worker.StartAsync(cancellation.Token), + Throws.Exception.With.Message.Contain("EndpointSettings is Copying (NotStarted)").And.Message.Not.Contain("KnownEndpoints is"), + "a worker let into a database whose copy stopped after its first category"); + Assert.That((await ReadCheckpoint(MigrationCategoryIds.EndpointSettings))?.State, Is.EqualTo(MigrationCategoryState.NotStarted)); + } + } + finally + { + await worker.DisposeAsync(); + } + } + + Task TargetEndpointSettingsCount() => + QueryTarget(dbContext => dbContext.EndpointSettings.CountAsync()); + + Task GetEndpointSettingsNames() => + QueryTarget(dbContext => dbContext.EndpointSettings.Select(settings => settings.Name).ToArrayAsync()); + + static void StoppedWhileStarting(WebApplicationBuilder builder) => + builder.Services.AddHostedService(provider => new StopsTheHostWhileStarting(provider.GetRequiredService())); + + static void CancelledWithoutAStop(WebApplicationBuilder builder) => + builder.Services.AddHostedService(_ => new CancelsWhileStarting()); + + sealed class CancelsWhileStarting : IHostedLifecycleService + { + public Task StartingAsync(CancellationToken cancellationToken = default) => throw new OperationCanceledException(); + + public Task StartAsync(CancellationToken cancellationToken = default) => Task.CompletedTask; + + public Task StartedAsync(CancellationToken cancellationToken = default) => Task.CompletedTask; + + public Task StoppingAsync(CancellationToken cancellationToken = default) => Task.CompletedTask; + + public Task StopAsync(CancellationToken cancellationToken = default) => Task.CompletedTask; + + public Task StoppedAsync(CancellationToken cancellationToken = default) => Task.CompletedTask; + } + + // What a service stop or the critical-error handler does to a host that is still starting. + sealed class StopsTheHostWhileStarting(IHostApplicationLifetime lifetime) : IHostedLifecycleService + { + public Task StartingAsync(CancellationToken cancellationToken = default) + { + lifetime.StopApplication(); + throw new OperationCanceledException(lifetime.ApplicationStopping); + } + + public Task StartAsync(CancellationToken cancellationToken = default) => Task.CompletedTask; + + public Task StartedAsync(CancellationToken cancellationToken = default) => Task.CompletedTask; + + public Task StoppingAsync(CancellationToken cancellationToken = default) => Task.CompletedTask; + + public Task StopAsync(CancellationToken cancellationToken = default) => Task.CompletedTask; + + public Task StoppedAsync(CancellationToken cancellationToken = default) => Task.CompletedTask; + } +} diff --git a/src/ServiceControl.Migration.AcceptanceTests/When_a_startup_check_fails.cs b/src/ServiceControl.Migration.AcceptanceTests/When_a_startup_check_fails.cs new file mode 100644 index 0000000000..86cd583c29 --- /dev/null +++ b/src/ServiceControl.Migration.AcceptanceTests/When_a_startup_check_fails.cs @@ -0,0 +1,266 @@ +namespace ServiceControl.Migration.AcceptanceTests; + +using System; +using System.IO; +using System.Linq; +using System.Net.Http; +using System.Security.Cryptography; +using System.Security.Cryptography.X509Certificates; +using System.Threading; +using System.Threading.Tasks; +using Microsoft.EntityFrameworkCore; +using Microsoft.EntityFrameworkCore.Infrastructure; +using Microsoft.EntityFrameworkCore.Migrations; +using NUnit.Framework; +using ServiceControl.Hosting.Commands; +using ServiceControl.Persistence.DataMigration; +using ServiceControl.Persistence.EFCore.Abstractions; +using ServiceControl.Persistence.EFCore.Entities; + +[TestFixture] +[NonParallelizable] +class When_a_startup_check_fails : MigrationAcceptanceTest +{ + const string HostOpenedSetting = "Migration/HostOpenedOnTarget"; + + [Test] + public async Task An_optional_category_window_that_is_not_a_time_span_is_refused() + { + SetSourceVariable("SERVICECONTROL_MIGRATION_EVENTLOGWINDOW", "a week"); + + var refusal = await RefusedStartup(); + + Assert.That(refusal.Message, Does.Contain("the optional category windows are valid").And.Contain("a week")); + } + + [Test] + public async Task A_retry_history_depth_that_empties_the_migrated_table_is_refused() + { + // Assigned rather than set as an environment variable: Settings reads RetryHistoryDepth once, in the constructor the fixture has already run. + Settings.RetryHistoryDepth = 0; + + var refusal = await RefusedStartup(); + + Assert.That(refusal.Message, Does.Contain("the retry history depth will not empty a migrated table") + .And.Contain("RetryHistoryDepth") + .And.Contain("HistoricRetryOperations")); + } + + [Test] + public async Task A_target_whose_schema_migrations_were_never_applied_is_refused() + { + await QueryTarget(async dbContext => + { + // EF's own history repository, because it is the only thing that knows where the history table is once the persister is given a schema. + var history = dbContext.GetService(); + var applied = await history.GetAppliedMigrationsAsync(); + + foreach (var row in applied) + { + await dbContext.Database.ExecuteSqlRawAsync(history.GetDeleteScript(row.MigrationId)); + } + + return applied.Count; + }); + + var refusal = await RefusedStartup(); + + Assert.That(refusal.Message, Does.Contain("the target schema is current").And.Contain("--setup")); + } + + [Test] + public async Task A_target_that_already_holds_data_is_refused() + { + await QueryTarget(dbContext => + { + dbContext.Settings.Add(new SettingEntity { Key = "NotificationEmails", Value = "{}" }); + return dbContext.SaveChangesAsync(); + }); + var settingsTable = await QueryTarget(dbContext => Task.FromResult(dbContext.Model.FindEntityType(typeof(SettingEntity))!.GetTableName()!)); + + var refusal = await RefusedStartup(); + + Assert.That(refusal.Message, Does.Contain("the migration target holds no ServiceControl data").And.Contain(settingsTable).And.Contain("--setup")); + } + + [Test] + public async Task Message_body_storage_that_cannot_be_written_to_is_refused() + { + var parentThatIsAFile = Path.Combine(Path.GetTempPath(), $"sc-not-a-directory-{Guid.NewGuid():n}"); + File.WriteAllText(parentThatIsAFile, string.Empty); + + try + { + ((FileSystemBodyStorageSettings)EFSettings.BodyStorage).StoragePath = Path.Combine(parentThatIsAFile, "bodies"); + + var refusal = await RefusedStartup(); + + Assert.That(refusal.Message, Does.Contain("message body storage is writable")); + } + finally + { + File.Delete(parentThatIsAFile); + } + } + + [Test] + public async Task A_client_certificate_that_has_expired_is_refused() + { + var notBefore = new DateTimeOffset(2020, 1, 1, 0, 0, 0, TimeSpan.Zero); + var notAfter = new DateTimeOffset(2021, 1, 1, 0, 0, 0, TimeSpan.Zero); + + SetSourceVariable("SERVICECONTROL_RAVENDB_CLIENTCERTIFICATEBASE64", ExpiredClientCertificate(notBefore, notAfter)); + + var refusal = await RefusedStartup(); + + Assert.That(refusal.Message, Does.Contain("the migration source opens") + .And.Contain($"{notBefore.UtcDateTime:u} to {notAfter.UtcDateTime:u}")); + } + + [Test] + public async Task A_throughput_database_that_does_not_exist_is_refused() + { + var missing = $"{Source.ThroughputDatabase}-does-not-exist"; + SetSourceVariable("LICENSINGCOMPONENT_RAVENDB_THROUGHPUTDATABASENAME", missing); + + var refusal = await RefusedStartup(); + + Assert.That(refusal.Message, Does.Contain("the migration source opens") + .And.Contain(missing) + .And.Contain("LicensingComponent/RavenDB/ThroughputDatabaseName")); + } + + // Without this check a build that copies the required set commits a customer to SQL before the rest of the migration exists. + [Test] + public async Task A_build_without_the_whole_migration_is_refused() + { + SetSourceVariable("SERVICECONTROL_MIGRATION_EVENTLOGWINDOW", "a week"); + + var refusal = await RefusedStartup(allowUnreleasedMigration: false); + + Assert.That(refusal.Message, Does.Contain("this build carries the whole migration") + .And.Not.Contain("the optional category windows are valid"), + "this refusal comes first, so no other check is consulted on an unreleased build"); + } + + [Test] + public async Task A_halted_required_category_stops_the_host_rather_than_opening_it() + { + await SeedSourceKnownEndpoints("Sales"); + await SeedSourceEndpointSettings(("Sales", true)); + + using var cancellation = new CancellationTokenSource(TimeSpan.FromMinutes(3)); + + var exception = Assert.ThrowsAsync(async () => + await RunCommand.Run(Settings, AllowingAnUnreleasedMigration(builder => builder.HaltTheEndpointSettingsCategory()), cancellation.Token)); + + Assert.Multiple(() => + { + Assert.That(exception.Message, Does.Contain("EndpointSettings").And.Contain("Halted")); + Assert.That(exception.Message, Does.Contain("nothing has been lost")); + Assert.That(exception.Message, Does.Contain("--migration-retry EndpointSettings").And.Contain("--migration-abandon EndpointSettings"), "a Failed category waits for the operator, so the refusal has to name what moves it on"); + Assert.That(async () => await HttpClient.GetAsync(EndpointSettingsUrl), Throws.InstanceOf(), "Kestrel bound its socket behind a halted required category"); + }); + } + + [Test] + public async Task A_source_that_will_not_open_once_the_required_copy_settled_still_lets_the_host_open() + { + await SeedSourceKnownEndpoints("Sales"); + await SeedSourceEndpointSettings(("Sales", true)); + await RunHostUntilTheApiAnswers(); + + var missing = $"{Source.PrimaryDatabase}-does-not-exist"; + SetSourceVariable("SERVICECONTROL_MIGRATION_EVENTLOGWINDOW", null); + SetSourceVariable("SERVICECONTROL_RAVENDB_DATABASENAME", missing); + + await RunHostUntilTheApiAnswers(); + + var endpointSettings = await GetEndpointSettings(); + var eventLog = await ReadCheckpoint(MigrationCategoryIds.EventLog); + var archived = await ReadCheckpoint(MigrationCategoryIds.ArchivedAndResolvedFailedMessages); + + using (Assert.EnterMultipleScope()) + { + Assert.That(endpointSettings.Select(setting => setting.Name), Does.Contain("Sales")); + Assert.That(eventLog?.State, Is.EqualTo(MigrationCategoryState.NotStarted)); + Assert.That(eventLog?.LastError, Does.Contain(missing).And.Not.Contain("will not start"), "status reads the outage from the row, on a start that did open"); + Assert.That(archived?.State, Is.EqualTo(MigrationCategoryState.NotStarted)); + Assert.That(archived?.LastError, Does.Contain(missing)); + } + } + + [Test] + public async Task A_source_that_will_not_open_before_the_required_copy_settled_still_refuses() + { + await SeedCheckpoint(MigrationCategoryIds.KnownEndpoints); + await SeedCheckpoint(MigrationCategoryIds.EndpointSettings, MigrationCategoryState.InProgress); + + var missing = $"{Source.PrimaryDatabase}-does-not-exist"; + SetSourceVariable("SERVICECONTROL_RAVENDB_DATABASENAME", missing); + + var refusal = await RefusedStartup(); + + using (Assert.EnterMultipleScope()) + { + Assert.That(refusal.Message, Does.Contain("the migration source opens").And.Contain(missing)); + Assert.That(refusal.Message, Does.Contain("--migration-abandon"), "a required category that never started has no exit once the source is gone, and the refusal has to say so"); + Assert.That(await ReadCheckpoint(MigrationCategoryIds.EventLog), Is.Null, "a refusal writes nothing"); + } + } + + [Test] + public async Task A_failed_optional_row_does_not_stop_the_host_opening() + { + await SeedSourceKnownEndpoints("Sales"); + await SeedSourceEndpointSettings(("Sales", true)); + await RunHostUntilTheApiAnswers(); + + await SeedCheckpoint(MigrationCategoryIds.EventLog, MigrationCategoryState.CompleteWithErrors); + SetSourceVariable("SERVICECONTROL_RAVENDB_DATABASENAME", $"{Source.PrimaryDatabase}-does-not-exist"); + + await RunHostUntilTheApiAnswers(); + + var eventLog = await ReadCheckpoint(MigrationCategoryIds.EventLog); + + using (Assert.EnterMultipleScope()) + { + Assert.That(eventLog?.State, Is.EqualTo(MigrationCategoryState.CompleteWithErrors)); + Assert.That(eventLog?.LastError, Is.Null, "the outage leaves a Failed row alone"); + } + } + + // Every refusal is asserted against a source that has rows to copy: an empty target proves nothing otherwise. + // Only the release gate's own test withholds the marker; every other refusal has to get past it. + async Task RefusedStartup(bool allowUnreleasedMigration = true) + { + await SeedSourceKnownEndpoints("Sales"); + await SeedSourceEndpointSettings(("Sales", true)); + + using var cancellation = new CancellationTokenSource(TimeSpan.FromMinutes(3)); + + var refusal = Assert.ThrowsAsync(async () => + await RunCommand.Run(Settings, allowUnreleasedMigration ? AllowingAnUnreleasedMigration() : null, cancellation.Token)); + + using (Assert.EnterMultipleScope()) + { + Assert.That(refusal.Message, Does.Contain("this start has copied nothing")); + Assert.That(await TargetEndpointSettingsCount(), Is.Zero, "a check that fires after rows have moved is worse than no check"); + Assert.That(await ReadSetting(HostOpenedSetting), Is.Null, "a refused startup has opened on nothing, so the abort is still free"); + } + + return refusal; + } + + Task TargetEndpointSettingsCount() => + QueryTarget(dbContext => dbContext.EndpointSettings.CountAsync()); + + static string ExpiredClientCertificate(DateTimeOffset notBefore, DateTimeOffset notAfter) + { + using var key = RSA.Create(2048); + var request = new CertificateRequest("CN=sc-migration-source-expired", key, HashAlgorithmName.SHA256, RSASignaturePadding.Pkcs1); + using var certificate = request.CreateSelfSigned(notBefore, notAfter); + + return Convert.ToBase64String(certificate.Export(X509ContentType.Pkcs12)); + } +} diff --git a/src/ServiceControl.Migration.AcceptanceTests/When_categories_change_between_restarts.cs b/src/ServiceControl.Migration.AcceptanceTests/When_categories_change_between_restarts.cs new file mode 100644 index 0000000000..07d17190d4 --- /dev/null +++ b/src/ServiceControl.Migration.AcceptanceTests/When_categories_change_between_restarts.cs @@ -0,0 +1,38 @@ +namespace ServiceControl.Migration.AcceptanceTests; + +using System.Threading.Tasks; +using NUnit.Framework; + +[TestFixture] +[NonParallelizable] +class When_categories_change_between_restarts : MigrationAcceptanceTest +{ + [Test] + public async Task Changing_the_windows_between_restarts_deletes_nothing() + { + await SeedSourceKnownEndpoints("Sales"); + await SeedSourceEndpointSettings(("Sales", true)); + + SetSourceVariable("SERVICECONTROL_MIGRATION_EVENTLOGWINDOW", "0"); + await RunHostUntilTheApiAnswers(); + Assert.That(await ReadCheckpoint("EventLog"), Is.Null, "a category no run has copied has no row, and no column anywhere says it was selected"); + var firstRun = await ReadCheckpoint("EndpointSettings"); + + SetSourceVariable("SERVICECONTROL_MIGRATION_EVENTLOGWINDOW", null); + await RunHostUntilTheApiAnswers(); + Assert.That(await ReadCheckpoint("EventLog"), Is.Null, "the required copy runs required categories only, so turning an optional one on starts nothing here"); + + SetSourceVariable("SERVICECONTROL_MIGRATION_EVENTLOGWINDOW", "0"); + await RunHostUntilTheApiAnswers(); + + var endpointSettings = await ReadCheckpoint("EndpointSettings"); + + Assert.Multiple(() => + { + Assert.That(endpointSettings.CopiedCount, Is.EqualTo(1), "a finished category must not be copied again when the selection changes around it"); + Assert.That(endpointSettings.Cursor, Is.Not.Null, "changing the selection must not delete or reset what an earlier run recorded"); + Assert.That(endpointSettings.Version, Is.EqualTo(firstRun.Version), "a row saved again, even unchanged, would have moved its version"); + Assert.That(endpointSettings.StartedAt, Is.EqualTo(firstRun.StartedAt), "a copy run again from scratch would have a new start"); + }); + } +} diff --git a/src/ServiceControl.Migration.AcceptanceTests/When_endpoint_settings_are_migrated.cs b/src/ServiceControl.Migration.AcceptanceTests/When_endpoint_settings_are_migrated.cs new file mode 100644 index 0000000000..29d5b8b432 --- /dev/null +++ b/src/ServiceControl.Migration.AcceptanceTests/When_endpoint_settings_are_migrated.cs @@ -0,0 +1,43 @@ +namespace ServiceControl.Migration.AcceptanceTests; + +using System; +using System.Linq; +using System.Threading.Tasks; +using Microsoft.Extensions.DependencyInjection; +using Microsoft.Extensions.Time.Testing; +using NUnit.Framework; + +[TestFixture] +[NonParallelizable] +class When_endpoint_settings_are_migrated : MigrationAcceptanceTest +{ + [Test] + public async Task The_source_server_starts_and_creates_both_databases() + { + var source = await MigrationSourceServer.CreateDatabases(); + + Assert.That(source.ServerUrl, Does.StartWith("http://localhost:")); + Assert.That(source.PrimaryDatabase, Is.Not.Empty); + Assert.That(source.ThroughputDatabase, Is.EqualTo($"{source.PrimaryDatabase}-throughput")); + } + + [Test] + public async Task They_are_readable_through_the_product_read_api() + { + await SeedSourceKnownEndpoints("Sales", "Billing"); + await SeedSourceEndpointSettings(("Sales", true), ("Billing", false), ("", true)); + + // A clock that never moves keeps HeartbeatEndpointSettingsSyncHostedService's twenty second + // delay from elapsing and rewriting these rows while the assertions read them. + await RunHostUntilTheApiAnswers(builder => builder.Services.AddSingleton(new FakeTimeProvider())); + + var settings = await GetEndpointSettings(); + + Assert.Multiple(() => + { + Assert.That(settings.Single(row => row.Name == "Sales").TrackInstances, Is.True); + Assert.That(settings.Single(row => row.Name == "Billing").TrackInstances, Is.False); + Assert.That(settings.Single(row => row.Name == string.Empty).TrackInstances, Is.True); + }); + } +} diff --git a/src/ServiceControl.Migration.AcceptanceTests/When_the_api_is_called_during_the_required_copy.cs b/src/ServiceControl.Migration.AcceptanceTests/When_the_api_is_called_during_the_required_copy.cs new file mode 100644 index 0000000000..a7be817c5d --- /dev/null +++ b/src/ServiceControl.Migration.AcceptanceTests/When_the_api_is_called_during_the_required_copy.cs @@ -0,0 +1,42 @@ +namespace ServiceControl.Migration.AcceptanceTests; + +using System; +using System.Linq; +using System.Net.Http; +using System.Threading; +using System.Threading.Tasks; +using NUnit.Framework; +using ServiceControl.Hosting.Commands; + +[TestFixture] +// Mandatory, not stylistic: this assembly is Parallelizable(ParallelScope.All) and these fixtures set +// process-global environment variables. One fixture added without it makes the whole suite intermittent. +[NonParallelizable] +class When_the_api_is_called_during_the_required_copy : MigrationAcceptanceTest +{ + [Test] + public async Task The_api_does_not_answer_until_the_required_copy_has_finished() + { + await SeedSourceKnownEndpoints("Sales", "Billing", "Shipping"); + await SeedSourceEndpointSettings(("Sales", true), ("Billing", true), ("Shipping", true)); + + var copyIsParked = new TaskCompletionSource(TaskCreationOptions.RunContinuationsAsynchronously); + var releaseTheCopy = new TaskCompletionSource(TaskCreationOptions.RunContinuationsAsynchronously); + + using var cancellation = new CancellationTokenSource(TimeSpan.FromMinutes(3)); + var host = RunCommand.Run(Settings, AllowingAnUnreleasedMigration(builder => builder.ParkFirstMigrationWrite(copyIsParked, releaseTheCopy.Task)), cancellation.Token); + + await copyIsParked.Task.WaitAsync(TimeSpan.FromMinutes(2)); + Assert.That(releaseTheCopy.Task.IsCompleted, Is.False, "the copy must still be parked for this assertion to mean anything"); + Assert.That(async () => await HttpClient.GetAsync(EndpointSettingsUrl), Throws.InstanceOf(), + "Kestrel had already bound its socket while the required copy was still running"); + + releaseTheCopy.SetResult(); + var settings = await WaitForEndpointSettingsResponse(TimeSpan.FromMinutes(2)); + + Assert.That(settings.Select(row => row.Name), Is.SupersetOf(new[] { "Sales", "Billing", "Shipping" })); + + await cancellation.CancelAsync(); + await host; + } +} diff --git a/src/ServiceControl.Migration.AcceptanceTests/When_the_host_opens_on_the_target.cs b/src/ServiceControl.Migration.AcceptanceTests/When_the_host_opens_on_the_target.cs new file mode 100644 index 0000000000..4c8d72c15a --- /dev/null +++ b/src/ServiceControl.Migration.AcceptanceTests/When_the_host_opens_on_the_target.cs @@ -0,0 +1,276 @@ +namespace ServiceControl.Migration.AcceptanceTests; + +using System; +using System.Linq; +using System.Threading; +using System.Threading.Tasks; +using Microsoft.AspNetCore.Builder; +using Microsoft.EntityFrameworkCore; +using Microsoft.EntityFrameworkCore.Infrastructure; +using Microsoft.EntityFrameworkCore.Storage; +using Microsoft.Extensions.DependencyInjection; +using Microsoft.Extensions.Hosting; +using NUnit.Framework; +using ServiceControl.Persistence.EFCore.Entities; +using ServiceControl.Persistence.DataMigration; +using ServiceControl.Hosting.Commands; + +[TestFixture] +// Mandatory, not stylistic: this assembly is Parallelizable(ParallelScope.All) and these fixtures set +// process-global environment variables. One fixture added without it makes the whole suite intermittent. +[NonParallelizable] +class When_the_host_opens_on_the_target : MigrationAcceptanceTest +{ + const string HostOpenedSetting = "Migration/HostOpenedOnTarget"; + const string CategoryFromANewerBuild = "CategoryFromANewerBuild"; + + [Test] + public async Task The_row_does_not_exist_while_the_required_copy_is_still_running() + { + await SeedSourceKnownEndpoints("Sales"); + await SeedSourceEndpointSettings(("Sales", true)); + + var copyIsParked = new TaskCompletionSource(TaskCreationOptions.RunContinuationsAsynchronously); + var releaseTheCopy = new TaskCompletionSource(TaskCreationOptions.RunContinuationsAsynchronously); + + using var cancellation = new CancellationTokenSource(TimeSpan.FromMinutes(3)); + var host = RunCommand.Run(Settings, AllowingAnUnreleasedMigration(builder => builder.ParkFirstMigrationWrite(copyIsParked, releaseTheCopy.Task)), cancellation.Token); + + await copyIsParked.Task.WaitAsync(TimeSpan.FromMinutes(2)); + Assert.That(await ReadSetting(HostOpenedSetting), Is.Null, "the abort is still free at this moment, and this row is what says it is not"); + + releaseTheCopy.SetResult(); + await WaitForEndpointSettingsResponse(TimeSpan.FromMinutes(2)); + + Assert.That(await ReadSetting(HostOpenedSetting), Is.Not.Null); + + await cancellation.CancelAsync(); + await host; + } + + [Test] + public async Task A_start_with_migration_off_does_not_stamp_the_target() + { + await SeedCheckpoint(MigrationCategoryIds.KnownEndpoints); + SetSourceVariable("SERVICECONTROL_MIGRATION_ENABLED", "false"); + var started = new TaskCompletionSource(TaskCreationOptions.RunContinuationsAsynchronously); + + await RunHostUntilTheApiAnswers(SignallingOnceStarted(started)); + await started.Task.WaitAsync(TimeSpan.FromMinutes(1)); + + Assert.That(await ReadSetting(HostOpenedSetting), Is.Null, "only a start with the migration on may write the marker"); + } + + [Test] + public async Task A_start_with_migration_off_does_not_read_the_checkpoint_table() + { + await DropTheCheckpointTable(); + SetSourceVariable("SERVICECONTROL_MIGRATION_ENABLED", "false"); + var started = new TaskCompletionSource(TaskCreationOptions.RunContinuationsAsynchronously); + + await RunHostUntilTheApiAnswers(SignallingOnceStarted(started)); + + Assert.That(async () => await started.Task.WaitAsync(TimeSpan.FromMinutes(1)), Throws.Nothing, "every hosted service, StartedAsync included, ran without touching the missing table"); + } + + [Test] + public async Task An_ingestion_only_host_is_refused_while_a_required_category_is_failed() + { + Settings.MaximumConcurrencyLevel = 1; + SetSourceVariable("SERVICECONTROL_MIGRATION_ENABLED", "false"); + await SeedRequiredCategoriesComplete(except: MigrationCategoryIds.MessageRedirects); + await SeedCheckpoint(MigrationCategoryIds.MessageRedirects, MigrationCategoryState.CompleteWithErrors); + + using var cancellation = new CancellationTokenSource(TimeSpan.FromMinutes(3)); + var app = ErrorIngestionOnlyCommand.BuildHost(Settings, AllowingAnUnreleasedMigration()); + + try + { + var refusal = Assert.ThrowsAsync(async () => await app.StartAsync(cancellation.Token)); + + Assert.That(refusal.Message, Does.Contain("MessageRedirects is Failed (CompleteWithErrors)") + .And.Contain("--migration-retry MessageRedirects") + .And.Contain("--migration-abandon MessageRedirects")); + } + finally + { + await app.DisposeAsync(); + } + } + + [Test] + public async Task An_ingestion_only_host_starts_while_an_optional_category_is_still_copying_or_failed() + { + Settings.MaximumConcurrencyLevel = 1; + SetSourceVariable("SERVICECONTROL_MIGRATION_ENABLED", "false"); + await SeedRequiredCategoriesComplete(); + await SeedCheckpoint(MigrationCategoryIds.EventLog, MigrationCategoryState.InProgress); + await SeedCheckpoint(MigrationCategoryIds.ArchivedAndResolvedFailedMessages, MigrationCategoryState.CompleteWithErrors); + + using var cancellation = new CancellationTokenSource(TimeSpan.FromMinutes(3)); + var app = ErrorIngestionOnlyCommand.BuildHost(Settings, AllowingAnUnreleasedMigration()); + + try + { + Assert.That(async () => await app.StartAsync(cancellation.Token), Throws.Nothing, "an optional category copies in the background after the host opens, so it never holds a worker back"); + await app.StopAsync(cancellation.Token); + } + finally + { + await app.DisposeAsync(); + } + } + + [Test] + public async Task An_ingestion_only_host_is_refused_over_a_category_this_build_does_not_know() + { + Settings.MaximumConcurrencyLevel = 1; + SetSourceVariable("SERVICECONTROL_MIGRATION_ENABLED", "false"); + await SeedRequiredCategoriesComplete(); + await SeedCheckpoint(CategoryFromANewerBuild, MigrationCategoryState.InProgress); + + using var cancellation = new CancellationTokenSource(TimeSpan.FromMinutes(3)); + var app = ErrorIngestionOnlyCommand.BuildHost(Settings, AllowingAnUnreleasedMigration()); + + try + { + var refusal = Assert.ThrowsAsync(async () => await app.StartAsync(cancellation.Token)); + + Assert.That(refusal.Message, Does.Contain($"{CategoryFromANewerBuild} is Copying (InProgress)"), "nobody here can say whether an id this build does not hold is optional, so it holds the worker back"); + } + finally + { + await app.DisposeAsync(); + } + } + + [Test] + public async Task Importing_failed_errors_is_refused_while_a_copy_is_unfinished() + { + SetSourceVariable("SERVICECONTROL_MIGRATION_ENABLED", "false"); + await SeedCheckpoint(MigrationCategoryIds.KnownEndpoints, MigrationCategoryState.InProgress); + + using var cancellation = new CancellationTokenSource(TimeSpan.FromMinutes(3)); + using var host = ImportFailedErrorsCommand.BuildHost(Settings); + + var refusal = Assert.ThrowsAsync(async () => await host.StartAsync(cancellation.Token)); + + Assert.That(refusal.Message, Does.Contain("--import-failed-errors will not start").And.Contain("KnownEndpoints is Copying (InProgress)")); + } + + [Test] + public async Task Importing_failed_errors_starts_while_an_optional_category_is_still_copying_or_failed() + { + SetSourceVariable("SERVICECONTROL_MIGRATION_ENABLED", "false"); + await SeedRequiredCategoriesComplete(); + await SeedCheckpoint(MigrationCategoryIds.EventLog, MigrationCategoryState.InProgress); + await SeedCheckpoint(MigrationCategoryIds.ArchivedAndResolvedFailedMessages, MigrationCategoryState.CompleteWithErrors); + + using var cancellation = new CancellationTokenSource(TimeSpan.FromMinutes(3)); + using var host = ImportFailedErrorsCommand.BuildHost(Settings); + + Assert.That(async () => await host.StartAsync(cancellation.Token), Throws.Nothing, "an optional category copies in the background after the host opens, so it never holds the import back"); + await host.StopAsync(cancellation.Token); + } + + // A customer whose migration never started must not be told they have passed the point of no return. + [Test] + public async Task The_row_does_not_exist_after_a_startup_check_refused() + { + // Seeded so the marker would be written if anything stamped it, which makes the refusal the only reason it is not. + await SeedCheckpoint(MigrationCategoryIds.KnownEndpoints); + SetSourceVariable("SERVICECONTROL_MIGRATION_EVENTLOGWINDOW", "NoSuchWindow"); + + using var cancellation = new CancellationTokenSource(TimeSpan.FromMinutes(3)); + + Assert.That(async () => await RunCommand.Run(Settings, AllowingAnUnreleasedMigration(), cancellation.Token), + Throws.Exception.With.Message.Contains("NoSuchWindow")); + + Assert.That(await ReadSetting(HostOpenedSetting), Is.Null, "a refused startup has opened on nothing"); + } + + // Ingestion-only hosts write failed messages into the same database a migration targets, so a rollback + // gate that had not seen them would discard everything they ingested. + [Test] + public async Task An_ingestion_only_host_records_the_row_too() + { + using var cancellation = new CancellationTokenSource(TimeSpan.FromMinutes(3)); + + // The ingestion-only host runs a real transport, which the migration fixture does not otherwise configure. + Settings.MaximumConcurrencyLevel = 1; + + await SeedCheckpoint(MigrationCategoryIds.KnownEndpoints); + + var app = ErrorIngestionOnlyCommand.BuildHost(Settings, AllowingAnUnreleasedMigration()); + + await app.StartAsync(cancellation.Token); + + try + { + Assert.That(await ReadSetting(HostOpenedSetting), Is.Not.Null, "only RunCommand used to record this, so a scaled-out ingestion host wrote to the target and left no trace of having opened on it"); + } + finally + { + await app.StopAsync(cancellation.Token); + await app.DisposeAsync(); + } + } + + + // A customer running on SQL who has never migrated must not be told the way back is gone. + [Test] + public async Task A_database_no_migration_has_touched_is_not_stamped() + { + Settings.MaximumConcurrencyLevel = 1; + + using var cancellation = new CancellationTokenSource(TimeSpan.FromMinutes(3)); + var app = ErrorIngestionOnlyCommand.BuildHost(Settings, AllowingAnUnreleasedMigration()); + + await app.StartAsync(cancellation.Token); + + try + { + Assert.That(await ReadSetting(HostOpenedSetting), Is.Null, "no checkpoint row exists, so no migration has ever run against this database"); + } + finally + { + await app.StopAsync(cancellation.Token); + await app.DisposeAsync(); + } + } + + async Task SeedRequiredCategoriesComplete(string except = null) + { + foreach (var category in MigrationCategoryRegistry.All.Where(category => category.Kind == MigrationCategoryKind.Required && category.Id != except)) + { + await SeedCheckpoint(category.Id); + } + } + + // ApplicationStarted fires only after every hosted service's StartedAsync, which is where the marker is written. + static Action SignallingOnceStarted(TaskCompletionSource started) => + builder => builder.Services.AddHostedService(provider => new SignalWhenStarted(provider.GetRequiredService(), started)); + + sealed class SignalWhenStarted(IHostApplicationLifetime lifetime, TaskCompletionSource started) : IHostedService + { + public Task StartAsync(CancellationToken cancellationToken = default) + { + lifetime.ApplicationStarted.Register(() => started.TrySetResult()); + return Task.CompletedTask; + } + + public Task StopAsync(CancellationToken cancellationToken = default) => Task.CompletedTask; + } + + Task DropTheCheckpointTable() => + QueryTarget(dbContext => + { + var table = dbContext.Model.FindEntityType(typeof(MigrationCheckpointEntity))!; + var name = dbContext.GetService().DelimitIdentifier(table.GetTableName()!, table.GetSchema()); + + // The name comes from the EF model and the provider delimits it, so no outside input reaches this SQL. +#pragma warning disable EF1003 + return dbContext.Database.ExecuteSqlRawAsync("DROP TABLE " + name); +#pragma warning restore EF1003 + }); +} diff --git a/src/ServiceControl.Migration.Tests/MigrationCategoryCoverageTests.cs b/src/ServiceControl.Migration.Tests/MigrationCategoryCoverageTests.cs new file mode 100644 index 0000000000..1bdc4ed89f --- /dev/null +++ b/src/ServiceControl.Migration.Tests/MigrationCategoryCoverageTests.cs @@ -0,0 +1,100 @@ +namespace ServiceControl.Migration.Tests; + +using System; +using System.Collections.Generic; +using System.IO; +using System.Linq; +using System.Threading.Tasks; +using Microsoft.Extensions.DependencyInjection; +using NUnit.Framework; +using ServiceBus.Management.Infrastructure.Settings; +using ServiceControl.Persistence; +using ServiceControl.Persistence.DataMigration; + +[TestFixture] +[NonParallelizable] +class MigrationCategoryCoverageTests +{ + Settings settings; + + [SetUp] + public void SetUp() + { + // Both persistence configurations read ErrorRetentionPeriod through the settings reader rather + // than off the Settings object, so the constructor argument below does not satisfy either of them. + SetVariable("SERVICECONTROL_ERRORRETENTIONPERIOD", "10.00:00:00"); + // Registering the SQL Server persistence never connects, so any connection string satisfies it. + SetVariable("SERVICECONTROL_DATABASE_CONNECTIONSTRING", "Server=localhost;Database=ServiceControl;Trusted_Connection=True;TrustServerCertificate=True"); + SetVariable("SERVICECONTROL_MESSAGEBODY_STORAGETYPE", "FileSystem"); + SetVariable("SERVICECONTROL_MESSAGEBODY_FILESYSTEM_STORAGEPATH", + Path.Combine(TestContext.CurrentContext.WorkDirectory, "Bodies", Guid.NewGuid().ToString("n"))); + + // SQL Server stands in for PostgreSQL too, because both register the same EFCoreMigrationTarget and its writer list does not depend on the provider. + settings = new Settings(persisterType: "SQLServer", forwardErrorMessages: false, errorRetentionPeriod: TimeSpan.FromDays(10)); + } + + [TearDown] + public void TearDown() + { + foreach (var name in variables) + { + Environment.SetEnvironmentVariable(name, null); + } + + variables.Clear(); + } + + // A reader with no writer, or the reverse, is a category the intersection would otherwise drop in silence. + [Test] + public async Task Every_reader_has_a_writer_and_the_reverse() + { + await using var source = PersistenceFactory.CreateMigrationSource(settings); + var sourceCategoryIds = source.SupportedCategoryIds; + + var services = new ServiceCollection(); + services.AddLogging(); + services.AddPersistence(settings); + await using var provider = services.BuildServiceProvider(); + var target = provider.GetRequiredService(); + + var readerOnly = sourceCategoryIds.Except(target.SupportedCategoryIds).ToArray(); + var writerOnly = target.SupportedCategoryIds.Except(sourceCategoryIds).ToArray(); + + using (Assert.EnterMultipleScope()) + { + Assert.That(readerOnly, Is.Empty, $"the reader claims a category the writer does not: {string.Join(", ", readerOnly)}"); + Assert.That(writerOnly, Is.Empty, $"the writer claims a category the reader does not: {string.Join(", ", writerOnly)}"); + } + } + + [Test] + public async Task Every_reader_and_its_writer_agree_on_the_document_type() + { + await using var source = PersistenceFactory.CreateMigrationSource(settings); + + var services = new ServiceCollection(); + services.AddLogging(); + services.AddPersistence(settings); + await using var provider = services.BuildServiceProvider(); + var target = provider.GetRequiredService(); + + var shared = source.DocumentTypes.Keys.Intersect(target.DocumentTypes.Keys).ToArray(); + Assert.That(shared, Is.Not.Empty, "the test proves nothing if no category has both a reader and a writer"); + + using (Assert.EnterMultipleScope()) + { + foreach (var categoryId in shared) + { + Assert.That(target.DocumentTypes[categoryId], Is.EqualTo(source.DocumentTypes[categoryId]), $"the {categoryId} reader yields a document type its writer does not take"); + } + } + } + + void SetVariable(string name, string value) + { + Environment.SetEnvironmentVariable(name, value); + variables.Add(name); + } + + readonly List variables = []; +} diff --git a/src/ServiceControl.Persistence.EFCore.PostgreSql/Migrations/20260916125150_AddMigrationCheckpoints.Designer.cs b/src/ServiceControl.Persistence.EFCore.PostgreSql/Migrations/20261007000244_AddMigrationCheckpoints.Designer.cs similarity index 97% rename from src/ServiceControl.Persistence.EFCore.PostgreSql/Migrations/20260916125150_AddMigrationCheckpoints.Designer.cs rename to src/ServiceControl.Persistence.EFCore.PostgreSql/Migrations/20261007000244_AddMigrationCheckpoints.Designer.cs index a8a3dd5b3b..f9b6c71345 100644 --- a/src/ServiceControl.Persistence.EFCore.PostgreSql/Migrations/20260916125150_AddMigrationCheckpoints.Designer.cs +++ b/src/ServiceControl.Persistence.EFCore.PostgreSql/Migrations/20261007000244_AddMigrationCheckpoints.Designer.cs @@ -13,7 +13,7 @@ namespace ServiceControl.Persistence.EFCore.PostgreSql.Migrations { [DbContext(typeof(PostgreSqlServiceControlDbContext))] - [Migration("20260916125150_AddMigrationCheckpoints")] + [Migration("20261007000244_AddMigrationCheckpoints")] partial class AddMigrationCheckpoints { /// @@ -374,7 +374,8 @@ protected override void BuildTargetModel(ModelBuilder modelBuilder) .HasColumnName("message_id"); b.Property("MessageType") - .HasColumnType("text") + .HasMaxLength(450) + .HasColumnType("character varying(450)") .HasColumnName("message_type"); b.Property("NumberOfProcessingAttempts") @@ -440,8 +441,16 @@ protected override void BuildTargetModel(ModelBuilder modelBuilder) b.HasIndex("TimeSent") .HasDatabaseName("ix_failed_messages_time_sent"); - b.HasIndex("Status", "LastModified") - .HasDatabaseName("ix_failed_messages_status_last_modified"); + b.HasIndex("Status", "LastTimeOfFailure") + .HasDatabaseName("ix_failed_messages_status_last_time_of_failure"); + + b.HasIndex("Status", "LastModified", "UniqueMessageId") + .HasDatabaseName("ix_failed_messages_status_last_modified_unique_message_id"); + + NpgsqlIndexBuilderExtensions.IncludeProperties(b.HasIndex("Status", "LastModified", "UniqueMessageId"), new[] { "FirstTimeOfFailure", "LastTimeOfFailure" }); + + b.HasIndex("Status", "MessageType", "UniqueMessageId") + .HasDatabaseName("ix_failed_messages_status_message_type_unique_message_id"); b.ToTable("failed_messages", (string)null); }); @@ -477,6 +486,8 @@ protected override void BuildTargetModel(ModelBuilder modelBuilder) b.HasIndex("Type", "GroupId") .HasDatabaseName("ix_failed_message_groups_type_group_id"); + NpgsqlIndexBuilderExtensions.IncludeProperties(b.HasIndex("Type", "GroupId"), new[] { "FailedMessageUniqueId", "Title" }); + b.ToTable("failed_message_groups", (string)null); }); @@ -754,6 +765,10 @@ protected override void BuildTargetModel(ModelBuilder modelBuilder) .HasColumnType("timestamp with time zone") .HasColumnName("started_at"); + b.Property("StartedWindowSeconds") + .HasColumnType("bigint") + .HasColumnName("started_window_seconds"); + b.Property("State") .HasColumnType("integer") .HasColumnName("state"); diff --git a/src/ServiceControl.Persistence.EFCore.PostgreSql/Migrations/20260916125150_AddMigrationCheckpoints.cs b/src/ServiceControl.Persistence.EFCore.PostgreSql/Migrations/20261007000244_AddMigrationCheckpoints.cs similarity index 94% rename from src/ServiceControl.Persistence.EFCore.PostgreSql/Migrations/20260916125150_AddMigrationCheckpoints.cs rename to src/ServiceControl.Persistence.EFCore.PostgreSql/Migrations/20261007000244_AddMigrationCheckpoints.cs index b2b97c730a..6eaf796b17 100644 --- a/src/ServiceControl.Persistence.EFCore.PostgreSql/Migrations/20260916125150_AddMigrationCheckpoints.cs +++ b/src/ServiceControl.Persistence.EFCore.PostgreSql/Migrations/20261007000244_AddMigrationCheckpoints.cs @@ -27,7 +27,8 @@ protected override void Up(MigrationBuilder migrationBuilder) settled_at = table.Column(type: "timestamp with time zone", nullable: true), last_error = table.Column(type: "text", nullable: true), already_present_count = table.Column(type: "bigint", nullable: false), - version = table.Column(type: "bigint", nullable: false) + version = table.Column(type: "bigint", nullable: false), + started_window_seconds = table.Column(type: "bigint", nullable: true) }, constraints: table => { diff --git a/src/ServiceControl.Persistence.EFCore.PostgreSql/Migrations/PostgreSqlServiceControlDbContextModelSnapshot.cs b/src/ServiceControl.Persistence.EFCore.PostgreSql/Migrations/PostgreSqlServiceControlDbContextModelSnapshot.cs index cfc74e1483..bdfd8aebaa 100644 --- a/src/ServiceControl.Persistence.EFCore.PostgreSql/Migrations/PostgreSqlServiceControlDbContextModelSnapshot.cs +++ b/src/ServiceControl.Persistence.EFCore.PostgreSql/Migrations/PostgreSqlServiceControlDbContextModelSnapshot.cs @@ -762,6 +762,10 @@ protected override void BuildModel(ModelBuilder modelBuilder) .HasColumnType("timestamp with time zone") .HasColumnName("started_at"); + b.Property("StartedWindowSeconds") + .HasColumnType("bigint") + .HasColumnName("started_window_seconds"); + b.Property("State") .HasColumnType("integer") .HasColumnName("state"); diff --git a/src/ServiceControl.Persistence.EFCore.PostgreSql/PostgreSqlMigrationSqlDialect.cs b/src/ServiceControl.Persistence.EFCore.PostgreSql/PostgreSqlMigrationSqlDialect.cs new file mode 100644 index 0000000000..fdac0674c0 --- /dev/null +++ b/src/ServiceControl.Persistence.EFCore.PostgreSql/PostgreSqlMigrationSqlDialect.cs @@ -0,0 +1,43 @@ +namespace ServiceControl.Persistence.EFCore.PostgreSql; + +using ServiceControl.Persistence.EFCore.DbContexts; +using ServiceControl.Persistence.EFCore.Infrastructure; + +/// +/// The migration's PostgreSQL statements. An insert that does nothing on a conflicting key is all it takes to +/// add only what is absent, and the keys it inserted come back from the same statement. +/// +class PostgreSqlMigrationSqlDialect : PostgreSqlDialect, IMigrationSqlDialect +{ + // Deterministic collations compare keys byte for byte, so there is no schema fact to read and none to name in the statement. + public Task Open(ServiceControlDbContext dbContext, CancellationToken cancellationToken = default) => Task.CompletedTask; + + public IEqualityComparer KeyComparer(Type entityType, string propertyName) => StringComparer.Ordinal; + + // A fixed chunk, because PostgreSQL's parameter limit is far away and the row width does not bring it closer. + public int RowsPerStatement(int parametersPerRow) => MaxRowsPerStatement; + + public async Task> InsertMissing(ServiceControlDbContext dbContext, IReadOnlyList rows, CancellationToken cancellationToken = default) where TEntity : class + { + var insert = MigrationInsert.For(dbContext); + var keyColumns = string.Join(", ", insert.KeyColumns); + var inserted = new List(rows.Count); + + foreach (var chunk in rows.Chunk(RowsPerStatement(insert.Columns.Count))) + { + inserted.AddRange(await insert.Execute( + dbContext, + $""" + INSERT INTO {insert.Table} ({string.Join(", ", insert.Columns)}) + VALUES + {ParameterRows(chunk.Length, insert.Columns.Count)} + ON CONFLICT ({keyColumns}) DO NOTHING + RETURNING {keyColumns} + """, + chunk, + cancellationToken)); + } + + return inserted; + } +} diff --git a/src/ServiceControl.Persistence.EFCore.PostgreSql/PostgreSqlPersistence.cs b/src/ServiceControl.Persistence.EFCore.PostgreSql/PostgreSqlPersistence.cs index 3cbe6c1549..898a4a6bd4 100644 --- a/src/ServiceControl.Persistence.EFCore.PostgreSql/PostgreSqlPersistence.cs +++ b/src/ServiceControl.Persistence.EFCore.PostgreSql/PostgreSqlPersistence.cs @@ -18,6 +18,7 @@ public void AddPersistence(IServiceCollection services) RegisterDataStores(services, settings); services.AddSingleton(); + services.AddSingleton(); services.AddSingleton(); services.AddSingleton(); services.AddSingleton(); diff --git a/src/ServiceControl.Persistence.EFCore.SqlServer/Migrations/20260916125144_AddMigrationCheckpoints.Designer.cs b/src/ServiceControl.Persistence.EFCore.SqlServer/Migrations/20261007000240_AddMigrationCheckpoints.Designer.cs similarity index 97% rename from src/ServiceControl.Persistence.EFCore.SqlServer/Migrations/20260916125144_AddMigrationCheckpoints.Designer.cs rename to src/ServiceControl.Persistence.EFCore.SqlServer/Migrations/20261007000240_AddMigrationCheckpoints.Designer.cs index 5dfcc173b9..2f862bcb80 100644 --- a/src/ServiceControl.Persistence.EFCore.SqlServer/Migrations/20260916125144_AddMigrationCheckpoints.Designer.cs +++ b/src/ServiceControl.Persistence.EFCore.SqlServer/Migrations/20261007000240_AddMigrationCheckpoints.Designer.cs @@ -12,7 +12,7 @@ namespace ServiceControl.Persistence.EFCore.SqlServer.Migrations { [DbContext(typeof(SqlServerServiceControlDbContext))] - [Migration("20260916125144_AddMigrationCheckpoints")] + [Migration("20261007000240_AddMigrationCheckpoints")] partial class AddMigrationCheckpoints { /// @@ -302,7 +302,8 @@ protected override void BuildTargetModel(ModelBuilder modelBuilder) .HasColumnType("nvarchar(450)"); b.Property("MessageType") - .HasColumnType("nvarchar(max)"); + .HasMaxLength(450) + .HasColumnType("nvarchar(450)"); b.Property("NumberOfProcessingAttempts") .HasColumnType("int"); @@ -353,6 +354,12 @@ protected override void BuildTargetModel(ModelBuilder modelBuilder) b.HasIndex("Status", "LastModified"); + SqlServerIndexBuilderExtensions.IncludeProperties(b.HasIndex("Status", "LastModified"), new[] { "FirstTimeOfFailure", "LastTimeOfFailure" }); + + b.HasIndex("Status", "LastTimeOfFailure"); + + b.HasIndex("Status", "MessageType", "UniqueMessageId"); + b.ToTable("FailedMessages"); }); @@ -380,6 +387,8 @@ protected override void BuildTargetModel(ModelBuilder modelBuilder) b.HasIndex("Type", "GroupId"); + SqlServerIndexBuilderExtensions.IncludeProperties(b.HasIndex("Type", "GroupId"), new[] { "Title" }); + b.ToTable("FailedMessageGroups"); }); @@ -602,6 +611,9 @@ protected override void BuildTargetModel(ModelBuilder modelBuilder) b.Property("StartedAt") .HasColumnType("datetime2"); + b.Property("StartedWindowSeconds") + .HasColumnType("bigint"); + b.Property("State") .HasColumnType("int"); diff --git a/src/ServiceControl.Persistence.EFCore.SqlServer/Migrations/20260916125144_AddMigrationCheckpoints.cs b/src/ServiceControl.Persistence.EFCore.SqlServer/Migrations/20261007000240_AddMigrationCheckpoints.cs similarity index 94% rename from src/ServiceControl.Persistence.EFCore.SqlServer/Migrations/20260916125144_AddMigrationCheckpoints.cs rename to src/ServiceControl.Persistence.EFCore.SqlServer/Migrations/20261007000240_AddMigrationCheckpoints.cs index b19e73d767..c3dab88c81 100644 --- a/src/ServiceControl.Persistence.EFCore.SqlServer/Migrations/20260916125144_AddMigrationCheckpoints.cs +++ b/src/ServiceControl.Persistence.EFCore.SqlServer/Migrations/20261007000240_AddMigrationCheckpoints.cs @@ -27,7 +27,8 @@ protected override void Up(MigrationBuilder migrationBuilder) SettledAt = table.Column(type: "datetime2", nullable: true), LastError = table.Column(type: "nvarchar(max)", nullable: true), AlreadyPresentCount = table.Column(type: "bigint", nullable: false), - Version = table.Column(type: "bigint", nullable: false) + Version = table.Column(type: "bigint", nullable: false), + StartedWindowSeconds = table.Column(type: "bigint", nullable: true) }, constraints: table => { diff --git a/src/ServiceControl.Persistence.EFCore.SqlServer/Migrations/SqlServerServiceControlDbContextModelSnapshot.cs b/src/ServiceControl.Persistence.EFCore.SqlServer/Migrations/SqlServerServiceControlDbContextModelSnapshot.cs index e8a1ad22cd..42b194da3a 100644 --- a/src/ServiceControl.Persistence.EFCore.SqlServer/Migrations/SqlServerServiceControlDbContextModelSnapshot.cs +++ b/src/ServiceControl.Persistence.EFCore.SqlServer/Migrations/SqlServerServiceControlDbContextModelSnapshot.cs @@ -608,6 +608,9 @@ protected override void BuildModel(ModelBuilder modelBuilder) b.Property("StartedAt") .HasColumnType("datetime2"); + b.Property("StartedWindowSeconds") + .HasColumnType("bigint"); + b.Property("State") .HasColumnType("int"); diff --git a/src/ServiceControl.Persistence.EFCore.SqlServer/SqlServerDialect.cs b/src/ServiceControl.Persistence.EFCore.SqlServer/SqlServerDialect.cs index 5835c1f39c..acae4e5f55 100644 --- a/src/ServiceControl.Persistence.EFCore.SqlServer/SqlServerDialect.cs +++ b/src/ServiceControl.Persistence.EFCore.SqlServer/SqlServerDialect.cs @@ -69,6 +69,11 @@ protected static string ParameterRows(int rowCount, int columnCount) return sql.ToString(); } + // The same rows as ParameterRows, each ending with its position in the batch, so a statement can keep the first of two duplicate keys. + protected static string ParameterRowsWithOrdinal(int rowCount, int columnCount) => + string.Join(",\n", Enumerable.Range(0, rowCount).Select(row => + $"({string.Join(", ", Enumerable.Range(0, columnCount).Select(column => $"@p{(row * columnCount) + column}"))}, {row})")); + protected static int MaxRowsPerStatement(int columns) => MaxParametersPerStatement / columns; // SQL Server's ceiling is 2100 (https://learn.microsoft.com/en-us/sql/sql-server/maximum-capacity-specifications-for-sql-server). diff --git a/src/ServiceControl.Persistence.EFCore.SqlServer/SqlServerMigrationSqlDialect.cs b/src/ServiceControl.Persistence.EFCore.SqlServer/SqlServerMigrationSqlDialect.cs new file mode 100644 index 0000000000..45d8ea8454 --- /dev/null +++ b/src/ServiceControl.Persistence.EFCore.SqlServer/SqlServerMigrationSqlDialect.cs @@ -0,0 +1,114 @@ +namespace ServiceControl.Persistence.EFCore.SqlServer; + +using Microsoft.EntityFrameworkCore; +using Microsoft.EntityFrameworkCore.Infrastructure; +using Microsoft.EntityFrameworkCore.Metadata; +using Microsoft.EntityFrameworkCore.Storage; +using ServiceControl.Persistence.EFCore.DbContexts; +using ServiceControl.Persistence.EFCore.Infrastructure; + +/// +/// The migration's SQL Server statements. Two keys that differ only in case can be one row here, because a +/// column's collation decides what counts as equal, so the statements compare keys under the collation the +/// schema gives them and the batch is de-duplicated the same way before it is offered to the table. +/// +class SqlServerMigrationSqlDialect : SqlServerDialect, IMigrationSqlDialect +{ + const int IgnoreCaseStyle = 1; + + OpenedKeys? OpenedState { get; set; } + + public async Task Open(ServiceControlDbContext dbContext, CancellationToken cancellationToken = default) + { + var sql = dbContext.GetService(); + var collations = new Dictionary<(string Table, string Column), string>(); + var comparers = new Dictionary<(Type EntityType, string Property), IEqualityComparer>(); + + foreach (var entityType in dbContext.Model.GetEntityTypes()) + { + if (entityType.GetTableName() is not { } tableName || entityType.FindPrimaryKey() is not { } key) + { + continue; + } + + var storeObject = StoreObjectIdentifier.Table(tableName, entityType.GetSchema()); + var table = sql.DelimitIdentifier(tableName, entityType.GetSchema()); + + foreach (var property in key.Properties.Where(property => property.ClrType == typeof(string))) + { + // A key property with no column of its own is refused by name when MigrationInsert builds the statement. + if (property.GetColumnName(storeObject) is not { } column) + { + continue; + } + + var collation = await dbContext.Database + .SqlQuery($""" + SELECT c.collation_name AS [Name], + CONVERT(int, COLLATIONPROPERTY(c.collation_name, 'ComparisonStyle')) AS [ComparisonStyle] + FROM sys.columns AS c + WHERE c.object_id = OBJECT_ID({table}) AND c.name = {column} AND c.collation_name IS NOT NULL + """) + .SingleOrDefaultAsync(cancellationToken) + ?? throw new InvalidOperationException($"The migration target found no collation for {tableName}.{column}, so the SQL Server schema is missing or not current. Run ServiceControl with --setup against this database first."); + + collations[(table, sql.DelimitIdentifier(column))] = collation.Name; + comparers[(entityType.ClrType, property.Name)] = (collation.ComparisonStyle & IgnoreCaseStyle) == 0 ? StringComparer.Ordinal : StringComparer.OrdinalIgnoreCase; + } + } + + OpenedState = new OpenedKeys(collations, comparers); + } + + public IEqualityComparer KeyComparer(Type entityType, string propertyName) => + Keys.Comparers.TryGetValue((entityType, propertyName), out var comparer) + ? comparer + : throw new ArgumentException($"{entityType.Name}.{propertyName} is not a string primary-key column, so it has no collation to compare under."); + + public int RowsPerStatement(int parametersPerRow) => MaxRowsPerStatement(parametersPerRow); + + public async Task> InsertMissing(ServiceControlDbContext dbContext, IReadOnlyList rows, CancellationToken cancellationToken = default) where TEntity : class + { + var insert = MigrationInsert.For(dbContext); + var columns = string.Join(", ", insert.Columns); + // The collation is an identifier sys.columns returned, and COLLATE takes no parameter. + var partitionBy = string.Join(", ", insert.KeyColumns.Select(column => + Keys.Collations.TryGetValue((insert.Table, column), out var collation) ? $"{column} COLLATE {collation}" : column)); + var inserted = new List(rows.Count); + + foreach (var chunk in rows.Chunk(RowsPerStatement(insert.Columns.Count))) + { + // HOLDLOCK, or two writers can both find a key missing and both insert it. + inserted.AddRange(await insert.Execute( + dbContext, + $""" + MERGE {insert.Table} WITH (HOLDLOCK) AS t + USING ( + SELECT {columns} + FROM ( + SELECT {columns}, ROW_NUMBER() OVER (PARTITION BY {partitionBy} ORDER BY [MigrationOrdinal]) AS [MigrationDuplicate] + FROM (VALUES + {ParameterRowsWithOrdinal(chunk.Length, insert.Columns.Count)} + ) AS v ({columns}, [MigrationOrdinal]) + ) AS deduplicated + WHERE [MigrationDuplicate] = 1 + ) AS s ({columns}) + ON {string.Join(" AND ", insert.KeyColumns.Select(column => $"t.{column} = s.{column}"))} + WHEN NOT MATCHED THEN INSERT ({columns}) VALUES ({string.Join(", ", insert.Columns.Select(column => $"s.{column}"))}) + OUTPUT {string.Join(", ", insert.KeyColumns.Select(column => $"inserted.{column}"))}; + """, + chunk, + cancellationToken)); + } + + return inserted; + } + + OpenedKeys Keys => OpenedState ?? throw new InvalidOperationException($"The SQL Server migration dialect is not open. Call {nameof(Open)} first."); + + sealed record OpenedKeys( + IReadOnlyDictionary<(string Table, string Column), string> Collations, + IReadOnlyDictionary<(Type EntityType, string Property), IEqualityComparer> Comparers); + + sealed record KeyColumnCollation(string Name, int ComparisonStyle); +} diff --git a/src/ServiceControl.Persistence.EFCore.SqlServer/SqlServerPersistence.cs b/src/ServiceControl.Persistence.EFCore.SqlServer/SqlServerPersistence.cs index dead62d3cc..b508cfd7d1 100644 --- a/src/ServiceControl.Persistence.EFCore.SqlServer/SqlServerPersistence.cs +++ b/src/ServiceControl.Persistence.EFCore.SqlServer/SqlServerPersistence.cs @@ -18,6 +18,7 @@ public void AddPersistence(IServiceCollection services) RegisterDataStores(services, settings); services.AddSingleton(); + services.AddSingleton(); services.AddSingleton(); services.AddSingleton(); services.AddSingleton(); diff --git a/src/ServiceControl.Persistence.EFCore/Abstractions/BasePersistence.cs b/src/ServiceControl.Persistence.EFCore/Abstractions/BasePersistence.cs index 6b61a58b40..46cb05d67e 100644 --- a/src/ServiceControl.Persistence.EFCore/Abstractions/BasePersistence.cs +++ b/src/ServiceControl.Persistence.EFCore/Abstractions/BasePersistence.cs @@ -7,6 +7,7 @@ namespace ServiceControl.Persistence.EFCore.Abstractions; using Particular.LicensingComponent.Persistence; using ServiceControl.CustomChecks; using ServiceControl.Operations.BodyStorage; +using ServiceControl.Persistence.EFCore.DataMigration; using ServiceControl.Persistence.EFCore.Implementation; using ServiceControl.Persistence.EFCore.Implementation.BodyStorage; using ServiceControl.Persistence.EFCore.Implementation.Recoverability; @@ -40,6 +41,8 @@ protected static void RegisterDataStores(IServiceCollection services, EFPersiste services.AddHostedService(p => p.GetRequiredService()); services.AddSingleton(); + services.AddSingleton(); + services.AddSingleton(); if (settings.RunRetentionSweep) { diff --git a/src/ServiceControl.Persistence.EFCore/DataMigration/BodyStorageIsWritableCheck.cs b/src/ServiceControl.Persistence.EFCore/DataMigration/BodyStorageIsWritableCheck.cs new file mode 100644 index 0000000000..5a5c116bdf --- /dev/null +++ b/src/ServiceControl.Persistence.EFCore/DataMigration/BodyStorageIsWritableCheck.cs @@ -0,0 +1,35 @@ +namespace ServiceControl.Persistence.EFCore.DataMigration; + +using System.Text; +using Infrastructure; +using ServiceControl.Persistence.DataMigration; + +/// +/// Writes, reads and deletes one probe body, so a body store that is misconfigured stops the copy before it +/// starts. Bodies can live on a file share, in Azure Blob Storage or in S3, and without this the first category +/// that carries bodies would find out part way through and skip every message it could not write. +/// +sealed class BodyStorageIsWritableCheck(IBodyStoragePersistence bodyStorage) : IMigrationStartupCheck +{ + const string ProbeBodyId = "migration-writable-probe"; + + public string Name => "message body storage is writable"; + + public async Task Run(CancellationToken cancellationToken = default) + { + await bodyStorage.WriteBody(ProbeBodyId, Encoding.UTF8.GetBytes(ProbeBodyId), "text/plain", cancellationToken); + + try + { + var probe = await bodyStorage.ReadBody(ProbeBodyId, cancellationToken) + ?? throw new Exception($"Message body storage accepted the probe body '{ProbeBodyId}' and then did not return it."); + + await probe.Stream.DisposeAsync(); + } + finally + { + // The probe is not migrated data, so it goes even when the read fails: a probe left behind is a body the source never had. + await bodyStorage.DeleteBodyIfExists(ProbeBodyId, cancellationToken); + } + } +} diff --git a/src/ServiceControl.Persistence.EFCore/DataMigration/EFCoreMigrationTarget.cs b/src/ServiceControl.Persistence.EFCore/DataMigration/EFCoreMigrationTarget.cs new file mode 100644 index 0000000000..09510292ff --- /dev/null +++ b/src/ServiceControl.Persistence.EFCore/DataMigration/EFCoreMigrationTarget.cs @@ -0,0 +1,113 @@ +namespace ServiceControl.Persistence.EFCore.DataMigration; + +using System.Collections.Frozen; +using Implementation; +using Infrastructure; +using Microsoft.EntityFrameworkCore; +using Microsoft.Extensions.DependencyInjection; +using Microsoft.Extensions.Logging; +using ServiceControl.Persistence.DataMigration; +using Writers; + +/// +/// Writes a migration into SQL Server or PostgreSQL. It holds one writer per category and knows nothing about +/// any of them beyond that, so a category this build cannot write is simply absent from +/// and never reaches the engine. +/// +sealed class EFCoreMigrationTarget(IServiceScopeFactory scopeFactory, IMigrationSqlDialect migrationDialect, ILogger logger) : DataStoreBase(scopeFactory), IMigrationTarget +{ + readonly FrozenDictionary writers = new IMigrationCategoryWriter[] + { + new KnownEndpointsWriter(migrationDialect), + new EndpointSettingsWriter(migrationDialect) + }.ToFrozenDictionary(writer => writer.CategoryId, StringComparer.Ordinal); + + public Task Open(CancellationToken cancellationToken = default) => + ExecuteWithDbContext((dbContext, token) => migrationDialect.Open(dbContext, token), cancellationToken); + + public Task BatchSizeFor(MigrationCategory category, CancellationToken cancellationToken = default) => + ExecuteWithDbContext((dbContext, _) => Task.FromResult(WriterFor(category).BatchSize(dbContext)), cancellationToken); + + public Task Write( + MigrationCategory category, + MigrationBatch batch, + MigrationCheckpoint checkpointToExtend, + CancellationToken cancellationToken = default) => + ExecuteWithDbContext(async (dbContext, writeToken) => + { + var prepared = await WriterFor(category).Prepare(dbContext, batch, writeToken); + + AccountForEveryRow(category, batch, prepared); + + var skipReasons = prepared.Skips.Count == 0 ? null : prepared.Skips.GroupBy(skip => skip.Reason).ToDictionary(group => group.Key, group => group.LongCount()); + + var (copied, alreadyPresent, saved) = await dbContext.Database.CreateExecutionStrategy().ExecuteAsync(async token => + { + // This block is retried, and an attempt that failed leaves the checkpoint it added still tracked, so the next attempt would insert it a second time. + dbContext.ChangeTracker.Clear(); + + await using var transaction = await dbContext.Database.BeginTransactionAsync(token); + + var inserted = await prepared.Insert(dbContext, token); + var present = AlreadyPresentIn(category, batch, inserted, prepared.Skips.Count); + var stored = await dbContext.UpsertCheckpoint( + checkpointToExtend.Extend(inserted, prepared.Skips.Count, present, skipReasons), + token); + await transaction.CommitAsync(token); + + return (inserted, present, stored); + }, writeToken); + + foreach (var (sourceId, reason, detail) in prepared.Skips) + { + logger.LogWarning("Skipped {SourceId} in category {CategoryId} as {SkipReason}: {Detail}", sourceId, category.Id, reason, detail); + } + + foreach (var (sourceId, detail) in prepared.Merges) + { + logger.LogWarning("Merged {SourceId} in category {CategoryId} into a row with the same key: {Detail}", sourceId, category.Id, detail); + } + + return new MigrationWriteResult( + saved, + copied, + prepared.Skips.Count, + [.. prepared.Skips.Select(skip => skip.SourceId)], + AlreadyPresent: alreadyPresent, + SkipReasons: skipReasons); + }, cancellationToken); + + // Checked before the insert, because the subtraction below can only catch a writer that over-reports. + internal static void AccountForEveryRow(MigrationCategory category, MigrationBatch batch, PreparedBatch prepared) + { + if (prepared.PreparedRowCount + prepared.Skips.Count != batch.Rows.Count) + { + throw new InvalidOperationException($"The {category.Id} writer prepared {prepared.PreparedRowCount} rows and skipped {prepared.Skips.Count} of the {batch.Rows.Count} rows in the batch. Every row must be one or the other, or the rows it dropped would be counted as rows the target already held."); + } + } + + // Without the throw, a miscount would quietly shrink the halt threshold's denominator instead of failing. + internal static int AlreadyPresentIn(MigrationCategory category, MigrationBatch batch, int copied, int skipped) + { + var alreadyPresent = batch.Rows.Count - copied - skipped; + + if (alreadyPresent < 0) + { + throw new InvalidOperationException($"The target copied {copied} and skipped {skipped} of the {batch.Rows.Count} rows in category {category.Id}, leaving {alreadyPresent} already present. Every row is copied, skipped or already present."); + } + + return alreadyPresent; + } + + public Task Count(MigrationCategory category, CancellationToken cancellationToken = default) => + ExecuteWithDbContext((dbContext, token) => WriterFor(category).Count(dbContext, token), cancellationToken); + + public IReadOnlyCollection SupportedCategoryIds => writers.Keys; + + public IReadOnlyDictionary DocumentTypes => writers.ToFrozenDictionary(pair => pair.Key, pair => pair.Value.DocumentType, StringComparer.Ordinal); + + IMigrationCategoryWriter WriterFor(MigrationCategory category) => + writers.TryGetValue(category.Id, out var writer) + ? writer + : throw new NotSupportedException($"The migration target cannot yet write the '{category.Id}' category."); +} diff --git a/src/ServiceControl.Persistence.EFCore/DataMigration/EFCoreMigrationTargetReadiness.cs b/src/ServiceControl.Persistence.EFCore/DataMigration/EFCoreMigrationTargetReadiness.cs new file mode 100644 index 0000000000..9988718ea8 --- /dev/null +++ b/src/ServiceControl.Persistence.EFCore/DataMigration/EFCoreMigrationTargetReadiness.cs @@ -0,0 +1,39 @@ +namespace ServiceControl.Persistence.EFCore.DataMigration; + +using Implementation; +using Infrastructure; +using Microsoft.Extensions.DependencyInjection; +using ServiceControl.Persistence.DataMigration; + +/// +/// What a SQL Server or PostgreSQL database has to be before a copy writes to it, and the stamp saying a host +/// has opened on it since. The stamp is a row in the settings table, so it survives a restart and belongs to +/// the database rather than to the instance. +/// +public class EFCoreMigrationTargetReadiness( + IServiceScopeFactory scopeFactory, + IBodyStoragePersistence bodyStorage, + TimeProvider timeProvider) : DataStoreBase(scopeFactory), IMigrationTargetReadiness +{ + // Cheapest first: the schema costs one query, and the body store costs a round trip to a file share or a + // cloud service. + public IReadOnlyList ContributedChecks() => + [ + new SchemaIsCurrentCheck(scopeFactory), + new TargetHoldsNoServiceControlDataCheck(scopeFactory), + new BodyStorageIsWritableCheck(bodyStorage) + ]; + + public Task RecordHostOpened(CancellationToken cancellationToken = default) => + ExecuteWithDbContext(async (dbContext, token) => + { + if (await dbContext.GetSetting(SettingKeys.MigrationHostOpenedOnTarget, token) is null) + { + await dbContext.StoreSetting(SettingKeys.MigrationHostOpenedOnTarget, timeProvider.GetUtcNow().UtcDateTime, token); + } + }, cancellationToken); + + public Task HasHostOpened(CancellationToken cancellationToken = default) => + ExecuteWithDbContext(async (dbContext, token) => + await dbContext.GetSetting(SettingKeys.MigrationHostOpenedOnTarget, token) is not null, cancellationToken); +} diff --git a/src/ServiceControl.Persistence.EFCore/DataMigration/IMigrationCategoryWriter.cs b/src/ServiceControl.Persistence.EFCore/DataMigration/IMigrationCategoryWriter.cs new file mode 100644 index 0000000000..3349a45a38 --- /dev/null +++ b/src/ServiceControl.Persistence.EFCore/DataMigration/IMigrationCategoryWriter.cs @@ -0,0 +1,38 @@ +namespace ServiceControl.Persistence.EFCore.DataMigration; + +using DbContexts; +using ServiceControl.Persistence.DataMigration; + +/// +/// Everything the target needs to know about one category: how big a batch is, how its documents +/// become rows, and how many rows the target holds for it. +/// +interface IMigrationCategoryWriter +{ + /// The category this writer handles, named as in . + string CategoryId { get; } + + /// + /// The type of takes. A row holding any other type + /// fails the batch with an . + /// + Type DocumentType { get; } + + /// + /// The most rows the target will take from the source in one batch. It comes from how many values each row + /// carries and how many parameters one statement can hold, so a wider table takes fewer rows. The context is read for its model only. + /// + int BatchSize(ServiceControlDbContext dbContext); + + /// + /// Turns one batch of source documents into the rows to insert and the rows to skip, writing nothing: the insert it returns runs later, inside the target's transaction. + /// Every batch row produces exactly one row or exactly one skip, never both and never neither, because the target works out how many were already there by subtracting both from the batch size. + /// + Task Prepare(ServiceControlDbContext dbContext, MigrationBatch batch, CancellationToken cancellationToken = default); + + /// + /// How many rows the target holds for this category. It counts only this category, even where two of them + /// share a table. + /// + Task Count(ServiceControlDbContext dbContext, CancellationToken cancellationToken = default); +} diff --git a/src/ServiceControl.Persistence.EFCore/DataMigration/MigrationCategoryWriter.cs b/src/ServiceControl.Persistence.EFCore/DataMigration/MigrationCategoryWriter.cs new file mode 100644 index 0000000000..4550acc0bb --- /dev/null +++ b/src/ServiceControl.Persistence.EFCore/DataMigration/MigrationCategoryWriter.cs @@ -0,0 +1,35 @@ +namespace ServiceControl.Persistence.EFCore.DataMigration; + +using DbContexts; +using ServiceControl.Persistence.DataMigration; + +/// +/// A writer that takes rows carrying a . It casts every row once, before +/// sees it, so a row from a reader that yields another type fails the batch by name. +/// +abstract class MigrationCategoryWriter : IMigrationCategoryWriter where TDocument : class +{ + public abstract string CategoryId { get; } + + public Type DocumentType => typeof(TDocument); + + public abstract int BatchSize(ServiceControlDbContext dbContext); + + public abstract Task Count(ServiceControlDbContext dbContext, CancellationToken cancellationToken = default); + + public Task Prepare(ServiceControlDbContext dbContext, MigrationBatch batch, CancellationToken cancellationToken = default) => + PrepareDocuments(dbContext, [.. batch.Rows.Select(row => (row, DocumentOf(row)))], cancellationToken); + + /// + /// Does what promises, for rows already cast. + /// + /// The context to read the target through. Nothing may be written through it here. + /// Every row in the batch, in source order, each with its document. + /// Cancels any read of the target. + /// One prepared row or one skip for every entry in . + protected abstract Task PrepareDocuments(ServiceControlDbContext dbContext, IReadOnlyList<(MigrationRow Row, TDocument Document)> documents, CancellationToken cancellationToken = default); + + TDocument DocumentOf(MigrationRow row) => + row.Document as TDocument + ?? throw new InvalidCastException($"The {CategoryId} writer takes {typeof(TDocument).FullName} documents, but row {row.SourceId} holds {row.Document?.GetType().FullName ?? "no document"}. The reader and writer for one category must agree on the document type."); +} diff --git a/src/ServiceControl.Persistence.EFCore/DataMigration/PreparedBatch.cs b/src/ServiceControl.Persistence.EFCore/DataMigration/PreparedBatch.cs new file mode 100644 index 0000000000..c4e3faebb6 --- /dev/null +++ b/src/ServiceControl.Persistence.EFCore/DataMigration/PreparedBatch.cs @@ -0,0 +1,18 @@ +namespace ServiceControl.Persistence.EFCore.DataMigration; + +using DbContexts; +using ServiceControl.Persistence.DataMigration; + +/// +/// One batch turned into rows, ready for the target to insert inside its own transaction. Nothing here has +/// touched the database yet. +/// +/// Runs the insert and returns how many rows it added. The target calls it inside the transaction that also saves the checkpoint. +/// How many rows the writer built, which is more than the insert adds when the table already holds some of their keys. +/// One entry per row the writer will not insert, with the reason and a detail line for the log. +/// One entry per prepared row whose key the target treats as the same key as another row's, either earlier in the batch or already in the table, so the insert keeps the other row and drops this one. Each carries a detail line naming both keys and the one kept, which the target logs as a warning after the commit. The dropped row is counted as already present. +sealed record PreparedBatch( + Func> Insert, + int PreparedRowCount, + IReadOnlyList<(string SourceId, MigrationSkipReason Reason, string Detail)> Skips, + IReadOnlyList<(string SourceId, string Detail)> Merges); diff --git a/src/ServiceControl.Persistence.EFCore/DataMigration/SchemaIsCurrentCheck.cs b/src/ServiceControl.Persistence.EFCore/DataMigration/SchemaIsCurrentCheck.cs new file mode 100644 index 0000000000..24f64d1de2 --- /dev/null +++ b/src/ServiceControl.Persistence.EFCore/DataMigration/SchemaIsCurrentCheck.cs @@ -0,0 +1,26 @@ +namespace ServiceControl.Persistence.EFCore.DataMigration; + +using Implementation; +using Microsoft.EntityFrameworkCore; +using Microsoft.Extensions.DependencyInjection; +using ServiceControl.Persistence.DataMigration; + +/// +/// Refuses a copy into a database whose schema is behind this build. The writers insert into tables by name, so +/// a missing schema migration turns into a failure part way through the first category instead of a refusal. +/// +sealed class SchemaIsCurrentCheck(IServiceScopeFactory scopeFactory) : DataStoreBase(scopeFactory), IMigrationStartupCheck +{ + public string Name => "the target schema is current"; + + public Task Run(CancellationToken cancellationToken = default) => + ExecuteWithDbContext(async (dbContext, token) => + { + string[] pending = [.. await dbContext.Database.GetPendingMigrationsAsync(token)]; + + if (pending.Length > 0) + { + throw new Exception($"The database is missing {pending.Length} schema migration(s): {string.Join(", ", pending)}. Run ServiceControl with --setup before starting a migration."); + } + }, cancellationToken); +} diff --git a/src/ServiceControl.Persistence.EFCore/DataMigration/TargetHoldsNoServiceControlDataCheck.cs b/src/ServiceControl.Persistence.EFCore/DataMigration/TargetHoldsNoServiceControlDataCheck.cs new file mode 100644 index 0000000000..b545a8bea6 --- /dev/null +++ b/src/ServiceControl.Persistence.EFCore/DataMigration/TargetHoldsNoServiceControlDataCheck.cs @@ -0,0 +1,74 @@ +namespace ServiceControl.Persistence.EFCore.DataMigration; + +using DbContexts; +using Implementation; +using Microsoft.EntityFrameworkCore; +using Microsoft.Extensions.DependencyInjection; +using ServiceControl.Persistence.DataMigration; + +/// +/// Refuses a copy into a database that already holds ServiceControl data. A copy into a database that ran on SQL +/// mixes the customer's rows with rows SQL already held, and a retry of a category could then delete rows the copy +/// never wrote. The check stands down once a row that is not known to be optional exists, because that means a copy +/// has started and what the tables hold is what it wrote. +/// +sealed class TargetHoldsNoServiceControlDataCheck(IServiceScopeFactory scopeFactory) : DataStoreBase(scopeFactory), IMigrationStartupCheck +{ + public string Name => "the migration target holds no ServiceControl data"; + + // One entry per mapped entity but the checkpoint: its CLR type, and a query asking whether its table holds a row. + internal static readonly IReadOnlyList<(Type Entity, Func> HasRows)> Tables = + [ + Entry(dbContext => dbContext.CustomChecks), + Entry(dbContext => dbContext.EndpointSettings), + Entry(dbContext => dbContext.KnownEndpoints), + Entry(dbContext => dbContext.FailedMessages), + Entry(dbContext => dbContext.FailedMessageGroups), + Entry(dbContext => dbContext.GroupComments), + Entry(dbContext => dbContext.MessageRedirects), + Entry(dbContext => dbContext.RetryBatches), + Entry(dbContext => dbContext.RetryBatchNowForwarding), + Entry(dbContext => dbContext.FailedMessageRetries), + Entry(dbContext => dbContext.FailedErrorImports), + Entry(dbContext => dbContext.Settings), + Entry(dbContext => dbContext.Subscriptions), + Entry(dbContext => dbContext.EventLogItems), + Entry(dbContext => dbContext.HistoricRetryOperations), + Entry(dbContext => dbContext.UnacknowledgedRetryOperations), + Entry(dbContext => dbContext.ArchiveOperations), + Entry(dbContext => dbContext.FailedMessageEdits), + Entry(dbContext => dbContext.LicensingEndpoints), + Entry(dbContext => dbContext.LicensingEndpointThroughput), + Entry(dbContext => dbContext.ExternalIntegrationDispatchRequests) + ]; + + // Both halves of an entry come from one DbSet, so the type a test compares and the table the query reads cannot differ. + static (Type Entity, Func> HasRows) Entry(Func> set) where T : class => + (typeof(T), (dbContext, cancellationToken) => set(dbContext).AnyAsync(cancellationToken)); + + public Task Run(CancellationToken cancellationToken = default) => + ExecuteWithDbContext(async (dbContext, token) => + { + var checkpointed = await dbContext.MigrationCheckpoints.Select(checkpoint => checkpoint.CategoryId).ToListAsync(token); + + if (checkpointed.Any(id => MigrationCategoryRegistry.Find(id)?.Kind != MigrationCategoryKind.Optional)) + { + return; + } + + var holdingRows = new List(); + + foreach (var (entity, hasRows) in Tables) + { + if (await hasRows(dbContext, token)) + { + holdingRows.Add(dbContext.Model.FindEntityType(entity)!.GetTableName()!); + } + } + + if (holdingRows.Count > 0) + { + throw new Exception($"The database already holds ServiceControl data in {string.Join(", ", holdingRows)}, so ServiceControl has already run on it, and a migration copies only into an empty database. Create a new database, run ServiceControl with --setup against it, and do not start ServiceControl on it before migrating; then start again with migration on."); + } + }, cancellationToken); +} diff --git a/src/ServiceControl.Persistence.EFCore/DataMigration/Writers/EndpointSettingsWriter.cs b/src/ServiceControl.Persistence.EFCore/DataMigration/Writers/EndpointSettingsWriter.cs new file mode 100644 index 0000000000..485518ca67 --- /dev/null +++ b/src/ServiceControl.Persistence.EFCore/DataMigration/Writers/EndpointSettingsWriter.cs @@ -0,0 +1,68 @@ +namespace ServiceControl.Persistence.EFCore.DataMigration.Writers; + +using DbContexts; +using Entities; +using Infrastructure; +using Microsoft.EntityFrameworkCore; +using ServiceControl.Persistence.DataMigration; + +/// +/// Writes every per-endpoint setting as stored. A setting for an endpoint the target does not know is copied too, +/// and the heartbeat settings sync removes it by its own rule once the instance opens. On a target whose name +/// column ignores case, two names that differ only in case are one row, and the writer reports each setting the +/// insert drops as a merge. +/// +sealed class EndpointSettingsWriter(IMigrationSqlDialect migrationDialect) : MigrationCategoryWriter +{ + public override string CategoryId => MigrationCategoryIds.EndpointSettings; + + public override int BatchSize(ServiceControlDbContext dbContext) => migrationDialect.RowsPerStatement(MigrationInsert.For(dbContext).Columns.Count); + + public override Task Count(ServiceControlDbContext dbContext, CancellationToken cancellationToken = default) => + dbContext.EndpointSettings.LongCountAsync(cancellationToken); + + protected override async Task PrepareDocuments(ServiceControlDbContext dbContext, IReadOnlyList<(MigrationRow Row, EndpointSettings Document)> documents, CancellationToken cancellationToken = default) + { + List rows = [.. documents.Select(document => new EndpointSettingsEntity { Name = document.Document.Name, TrackInstances = document.Document.TrackInstances })]; + List rowSourceIds = [.. documents.Select(document => document.Row.SourceId)]; + + // The source yields document-id order, and the statement keeps the first of any keys the database treats as one. + return new PreparedBatch( + async (context, token) => (await migrationDialect.InsertMissing(context, rows, token)).Count, + rows.Count, + [], + await MergesIn(dbContext, rows, rowSourceIds, cancellationToken)); + } + + // Mirrors the insert's de-duplication, so an operator can see which settings a case-insensitive name column dropped. + async Task> MergesIn(ServiceControlDbContext dbContext, List rows, List rowSourceIds, CancellationToken cancellationToken) + { + var sameName = migrationDialect.KeyComparer(typeof(EndpointSettingsEntity), nameof(EndpointSettingsEntity.Name)); + string[] rowNames = [.. rows.Select(row => row.Name)]; + + // The database compares under the column's collation, so this also returns a stored name that differs only in case. + var keptNames = await dbContext.EndpointSettings.AsNoTracking() + .Select(stored => stored.Name) + .Where(name => rowNames.Contains(name)) + .ToListAsync(cancellationToken); + + var merges = new List<(string SourceId, string Detail)>(); + + for (var index = 0; index < rows.Count; index++) + { + var name = rows[index].Name; + var keptName = keptNames.FirstOrDefault(kept => sameName.Equals(kept, name)); + + if (keptName is null) + { + keptNames.Add(name); + } + else if (!string.Equals(keptName, name, StringComparison.Ordinal)) + { + merges.Add((rowSourceIds[index], $"the name column treats '{name}' and '{keptName}' as one endpoint, so it kept '{keptName}' and dropped the settings of '{name}'")); + } + } + + return merges; + } +} diff --git a/src/ServiceControl.Persistence.EFCore/DataMigration/Writers/KnownEndpointsWriter.cs b/src/ServiceControl.Persistence.EFCore/DataMigration/Writers/KnownEndpointsWriter.cs new file mode 100644 index 0000000000..a0ec9ff796 --- /dev/null +++ b/src/ServiceControl.Persistence.EFCore/DataMigration/Writers/KnownEndpointsWriter.cs @@ -0,0 +1,52 @@ +namespace ServiceControl.Persistence.EFCore.DataMigration.Writers; + +using DbContexts; +using Entities; +using Infrastructure; +using Microsoft.EntityFrameworkCore; +using ServiceControl.Persistence.DataMigration; + +/// +/// Writes the endpoints ServiceControl has heard from. The row's key is worked out from the endpoint's own +/// details rather than carried over, so the same endpoint lands on the same row whichever side wrote it first. +/// +sealed class KnownEndpointsWriter(IMigrationSqlDialect migrationDialect) : MigrationCategoryWriter +{ + public override string CategoryId => MigrationCategoryIds.KnownEndpoints; + + public override int BatchSize(ServiceControlDbContext dbContext) => migrationDialect.RowsPerStatement(MigrationInsert.For(dbContext).Columns.Count); + + public override Task Count(ServiceControlDbContext dbContext, CancellationToken cancellationToken = default) => + dbContext.KnownEndpoints.LongCountAsync(cancellationToken); + + protected override Task PrepareDocuments(ServiceControlDbContext dbContext, IReadOnlyList<(MigrationRow Row, KnownEndpoint Document)> documents, CancellationToken cancellationToken = default) + { + var rows = new List(documents.Count); + var skips = new List<(string SourceId, MigrationSkipReason Reason, string Detail)>(); + + foreach (var (row, endpoint) in documents) + { + if (endpoint.EndpointDetails?.Name is null || endpoint.EndpointDetails.Host is null) + { + var column = endpoint.EndpointDetails?.Name is null ? nameof(KnownEndpointEntity.Name) : nameof(KnownEndpointEntity.Host); + skips.Add((row.SourceId, MigrationSkipReason.RequiredValueMissing, $"KnownEndpoints.{column} is NOT NULL and the document has no value for it")); + continue; + } + + rows.Add(new KnownEndpointEntity + { + Id = endpoint.EndpointDetails.GetDeterministicId(), + Name = endpoint.EndpointDetails.Name, + HostId = endpoint.EndpointDetails.HostId, + Host = endpoint.EndpointDetails.Host, + Monitored = endpoint.Monitored + }); + } + + return Task.FromResult(new PreparedBatch( + async (context, token) => (await migrationDialect.InsertMissing(context, rows, token)).Count, + rows.Count, + skips, + Merges: [])); + } +} diff --git a/src/ServiceControl.Persistence.EFCore/Entities/MigrationCheckpointEntity.cs b/src/ServiceControl.Persistence.EFCore/Entities/MigrationCheckpointEntity.cs index 9f82eb4335..2f43e79e78 100644 --- a/src/ServiceControl.Persistence.EFCore/Entities/MigrationCheckpointEntity.cs +++ b/src/ServiceControl.Persistence.EFCore/Entities/MigrationCheckpointEntity.cs @@ -17,10 +17,11 @@ public class MigrationCheckpointEntity public string? LastError { get; set; } public long AlreadyPresentCount { get; set; } public long Version { get; set; } + public long? StartedWindowSeconds { get; set; } internal MigrationCheckpoint ToCheckpoint() => new( CategoryId, State, Cursor, CopiedCount, SkippedCount, SourceTotal, SkipReasons, StartedAt, LastProgressAt, SettledAt, LastError, - AlreadyPresentCount, Version); + AlreadyPresentCount, Version, StartedWindowSeconds); } \ No newline at end of file diff --git a/src/ServiceControl.Persistence.EFCore/Infrastructure/IMigrationSqlDialect.cs b/src/ServiceControl.Persistence.EFCore/Infrastructure/IMigrationSqlDialect.cs new file mode 100644 index 0000000000..864562c443 --- /dev/null +++ b/src/ServiceControl.Persistence.EFCore/Infrastructure/IMigrationSqlDialect.cs @@ -0,0 +1,31 @@ +namespace ServiceControl.Persistence.EFCore.Infrastructure; + +using ServiceControl.Persistence.EFCore.DbContexts; + +/// +/// The migration's provider-specific SQL. Statements run inside the caller's transaction and insert only what is absent. +/// +public interface IMigrationSqlDialect +{ + /// + /// Reads, once, whatever the statements need from the schema. On SQL Server that is the collation of every + /// string column of every primary key, composite keys included. + /// + Task Open(ServiceControlDbContext dbContext, CancellationToken cancellationToken = default); + + /// + /// How the database compares one string key column, so a caller can work out which of its own rows the insert will treat as one. The insert compares in the database itself; this only predicts what it will find. + /// Call first, and name a property that is a string column of the entity's primary key. + /// + IEqualityComparer KeyComparer(Type entityType, string propertyName); + + /// + /// How many rows the provider will take in one statement, for rows carrying this many values each. + /// + int RowsPerStatement(int parametersPerRow); + + /// + /// Inserts the rows whose key the table does not hold, merges any whose keys the database treats as one, and returns the rows it inserted. + /// + Task> InsertMissing(ServiceControlDbContext dbContext, IReadOnlyList rows, CancellationToken cancellationToken = default) where TEntity : class; +} diff --git a/src/ServiceControl.Persistence.EFCore/Infrastructure/MigrationInsert.cs b/src/ServiceControl.Persistence.EFCore/Infrastructure/MigrationInsert.cs new file mode 100644 index 0000000000..37ac476786 --- /dev/null +++ b/src/ServiceControl.Persistence.EFCore/Infrastructure/MigrationInsert.cs @@ -0,0 +1,141 @@ +namespace ServiceControl.Persistence.EFCore.Infrastructure; + +using Microsoft.EntityFrameworkCore; +using Microsoft.EntityFrameworkCore.Infrastructure; +using Microsoft.EntityFrameworkCore.Metadata; +using Microsoft.EntityFrameworkCore.Storage; +using ServiceControl.Persistence.EFCore.DbContexts; + +/// +/// The table, columns and primary key maps to, and the run of one provider's insert statement over a chunk of those rows. +/// +public sealed class MigrationInsert where TEntity : class +{ + static readonly ExactKey KeyEquality = new(); + + readonly IReadOnlyList properties; + readonly IReadOnlyList keyProperties; + + MigrationInsert(string table, IReadOnlyList properties, IReadOnlyList keyProperties, Func columnName) + { + Table = table; + this.properties = properties; + this.keyProperties = keyProperties; + Columns = [.. properties.Select(columnName)]; + KeyColumns = [.. keyProperties.Select(columnName)]; + } + + /// The table name, quoted for this provider and carrying its schema. + public string Table { get; } + + /// Every column the insert writes, quoted, in the order the parameters of one row follow. + public IReadOnlyList Columns { get; } + + /// The primary key columns, quoted, in the order the statement has to return them. + public IReadOnlyList KeyColumns { get; } + + /// + /// Reads the table, columns and key out of the EF Core model for this entity. + /// + /// The entity is not mapped, has no primary key, or has a column this insert cannot write, such as one the database generates. Rows like that go through EF in the caller's context instead. + public static MigrationInsert For(ServiceControlDbContext dbContext) + { + var entityType = dbContext.Model.FindEntityType(typeof(TEntity)) + ?? throw new InvalidOperationException($"{typeof(TEntity).Name} is not an entity in the EF Core model, so the migration cannot tell which table it lands in."); + var tableName = entityType.GetTableName() + ?? throw new InvalidOperationException($"{typeof(TEntity).Name} is not mapped to a table."); + var key = entityType.FindPrimaryKey() + ?? throw new InvalidOperationException($"{typeof(TEntity).Name} has no primary key, so there is nothing for a row to be absent by."); + IProperty[] properties = [.. entityType.GetProperties()]; + + if (properties.Any(property => property.ValueGenerated != ValueGenerated.Never) + || entityType.GetComplexProperties().Any() + || entityType.GetNavigations().Any(navigation => navigation.TargetEntityType.IsOwned())) + { + throw new InvalidOperationException($"{typeof(TEntity).Name} has a store-generated, complex or owned property, whose columns InsertMissing does not write. Add these rows through EF in the caller's context instead."); + } + + var storeObject = StoreObjectIdentifier.Table(tableName, entityType.GetSchema()); + var sql = dbContext.GetService(); + + return new MigrationInsert( + sql.DelimitIdentifier(tableName, entityType.GetSchema()), + properties, + key.Properties, + property => sql.DelimitIdentifier(property.GetColumnName(storeObject) + ?? throw new InvalidOperationException($"{typeof(TEntity).Name}.{property.Name} has no column in {tableName}."))); + } + + /// + /// Runs one insert statement over a chunk of rows and returns the rows of that chunk it inserted. + /// The statement reads parameters named @p0 upward, row after row in order, and must return the of each row it inserted, in that order. + /// It runs on the caller's open transaction, and throws when there is none or when a returned key matches no row of the chunk. + /// + public async Task> Execute(ServiceControlDbContext dbContext, string sql, TEntity[] chunk, CancellationToken cancellationToken = default) + { + await using var command = dbContext.Database.GetDbConnection().CreateCommand(); + command.Transaction = (dbContext.Database.CurrentTransaction + ?? throw new InvalidOperationException("InsertMissing must run inside the caller's transaction, so its rows commit with the checkpoint.")).GetDbTransaction(); + command.CommandText = sql; + // A raw command starts at the provider's own 30 seconds, where SaveChangesAsync took the configured value. + command.CommandTimeout = dbContext.Database.GetCommandTimeout() ?? command.CommandTimeout; + + var index = 0; + foreach (var row in chunk) + { + foreach (var property in properties) + { + // The mapping applies the value converter and the provider type, as EF's own inserts do. + command.Parameters.Add(property.GetRelationalTypeMapping().CreateParameter(command, $"@p{index++}", property.GetGetter().GetClrValue(row), property.IsNullable)); + } + } + + var rowsByKey = new Dictionary(KeyEquality); + foreach (var row in chunk) + { + // A key repeated exactly is one row to every database, so the first row given carries it. + rowsByKey.TryAdd(KeyOf(row), row); + } + + var inserted = new List(chunk.Length); + + await using var reader = await command.ExecuteReaderAsync(cancellationToken); + while (await reader.ReadAsync(cancellationToken)) + { + object?[] key = [.. keyProperties.Select((property, column) => FromProvider(property, reader.GetValue(column)))]; + + // A key can come back as a different CLR type than it went in as, a date column read back as + // DateTime say, and a plain lookup would report that as a bare KeyNotFoundException. + inserted.Add(rowsByKey.TryGetValue(key, out var row) + ? row + : throw new InvalidOperationException($"The {Table} insert returned the key [{string.Join(", ", key.Select(Describe))}], which matches no row in the batch that produced it. Compare it against the key types the batch carried: [{string.Join(", ", keyProperties.Select(property => property.ClrType.Name))}].")); + } + + return inserted; + } + + object?[] KeyOf(TEntity row) => [.. keyProperties.Select(property => property.GetGetter().GetClrValue(row))]; + + static string Describe(object? value) => value is null ? "null" : $"{value} ({value.GetType().Name})"; + + static object? FromProvider(IProperty property, object value) => + property.GetRelationalTypeMapping().Converter is { } converter ? converter.ConvertFromProvider(value) : value; + + // Exact, because the database decided which keys were one before the statement ran, and OUTPUT and RETURNING echo the value inserted. + sealed class ExactKey : IEqualityComparer + { + public bool Equals(object?[]? x, object?[]? y) => x!.SequenceEqual(y!); + + public int GetHashCode(object?[] key) + { + var hash = new HashCode(); + + foreach (var value in key) + { + hash.Add(value); + } + + return hash.ToHashCode(); + } + } +} diff --git a/src/ServiceControl.Persistence.EFCore/Infrastructure/SettingKeys.cs b/src/ServiceControl.Persistence.EFCore/Infrastructure/SettingKeys.cs index 7a8e403170..7da4cc4bfe 100644 --- a/src/ServiceControl.Persistence.EFCore/Infrastructure/SettingKeys.cs +++ b/src/ServiceControl.Persistence.EFCore/Infrastructure/SettingKeys.cs @@ -9,4 +9,7 @@ static class SettingKeys public const string ReportMasks = "ReportMasks"; public const string LicensedEndpointDetails = "LicensedEndpointDetails"; public const string NotificationEmails = "NotificationEmails"; + // Written the first time a host starts on a database a migration has already written to. Once it is set, + // going back to RavenDB loses everything ServiceControl has written here since. + public const string MigrationHostOpenedOnTarget = "Migration/HostOpenedOnTarget"; } diff --git a/src/ServiceControl.Persistence.EFCore/MigrationCheckpointExtensions.cs b/src/ServiceControl.Persistence.EFCore/MigrationCheckpointExtensions.cs index e497fa1794..db12c9b895 100644 --- a/src/ServiceControl.Persistence.EFCore/MigrationCheckpointExtensions.cs +++ b/src/ServiceControl.Persistence.EFCore/MigrationCheckpointExtensions.cs @@ -91,5 +91,6 @@ static void Apply(MigrationCheckpoint checkpoint, MigrationCheckpointEntity enti entity.SettledAt = checkpoint.SettledAt; entity.LastError = checkpoint.LastError; entity.AlreadyPresentCount = checkpoint.AlreadyPresentCount; + entity.StartedWindowSeconds = checkpoint.StartedWindowSeconds; } } diff --git a/src/ServiceControl.Persistence.RavenDB/DataMigration/IMigrationCategoryReader.cs b/src/ServiceControl.Persistence.RavenDB/DataMigration/IMigrationCategoryReader.cs new file mode 100644 index 0000000000..8c05a27f97 --- /dev/null +++ b/src/ServiceControl.Persistence.RavenDB/DataMigration/IMigrationCategoryReader.cs @@ -0,0 +1,25 @@ +#nullable enable + +namespace ServiceControl.Persistence.RavenDB.DataMigration; + +using System; +using System.Collections.Generic; +using System.Threading; +using ServiceControl.Persistence.DataMigration; + +/// +/// Which documents one category is made of, and how they become rows. Read keeps the contract +/// sets out, for this one category. +/// +interface IMigrationCategoryReader +{ + /// The category this reader handles, named as in . + string CategoryId { get; } + + /// + /// The type of in every row yields. + /// + Type DocumentType { get; } + + IAsyncEnumerable Read(string? resumeAfter, int batchSize, CancellationToken cancellationToken = default); +} diff --git a/src/ServiceControl.Persistence.RavenDB/DataMigration/MigrationCategoryReader.cs b/src/ServiceControl.Persistence.RavenDB/DataMigration/MigrationCategoryReader.cs new file mode 100644 index 0000000000..590e534e18 --- /dev/null +++ b/src/ServiceControl.Persistence.RavenDB/DataMigration/MigrationCategoryReader.cs @@ -0,0 +1,40 @@ +#nullable enable + +namespace ServiceControl.Persistence.RavenDB.DataMigration; + +using System; +using System.Collections.Generic; +using System.Threading; +using ServiceControl.Persistence.DataMigration; + +/// +/// A reader whose every row carries a , which is the type its category's writer +/// must take. Reading through ties the documents RavenDB streams to that type. +/// +abstract class MigrationCategoryReader(RavenReadOnlySourceLifecycle lifecycle) : IMigrationCategoryReader where TDocument : class +{ + /// + /// The source connection, for a reader that streams or opens sessions itself instead of using . + /// + protected RavenReadOnlySourceLifecycle Lifecycle { get; } = lifecycle; + + public abstract string CategoryId { get; } + + public Type DocumentType => typeof(TDocument); + + public abstract IAsyncEnumerable Read(string? resumeAfter, int batchSize, CancellationToken cancellationToken = default); + + /// + /// Streams the primary database's documents whose id starts with , each one whole as a row. + /// + protected IAsyncEnumerable WholeDocuments(string prefix, string? resumeAfter, int batchSize, CancellationToken cancellationToken = default) => + RavenDocumentStream.ByPrefix( + Lifecycle, + CategoryId, + Lifecycle.Settings.DatabaseName, + prefix, + resumeAfter, + batchSize, + RavenDocumentStream.WholeDocument, + cancellationToken); +} diff --git a/src/ServiceControl.Persistence.RavenDB/DataMigration/RavenDocumentStream.cs b/src/ServiceControl.Persistence.RavenDB/DataMigration/RavenDocumentStream.cs new file mode 100644 index 0000000000..892641c1f5 --- /dev/null +++ b/src/ServiceControl.Persistence.RavenDB/DataMigration/RavenDocumentStream.cs @@ -0,0 +1,82 @@ +#nullable enable + +namespace ServiceControl.Persistence.RavenDB.DataMigration; + +using System; +using System.Collections.Generic; +using System.Collections.ObjectModel; +using System.Runtime.CompilerServices; +using System.Threading; +using Raven.Client.Documents.Commands; +using ServiceControl.Persistence.DataMigration; + +/// +/// The one way every category reads RavenDB: stream the documents whose id starts with a prefix, in id order, +/// and hand them back in batches. Id order is what makes the cursor work, because a resume asks RavenDB to +/// start after an id rather than to skip a count. +/// +static class RavenDocumentStream +{ + /// + /// Streams one collection and yields it in batches of at most . + /// + /// The document id prefix that selects the collection, such as "KnownEndpoints/". + /// The cursor a previous run saved, or null to start at the beginning. + /// Turns one document into a row, or into null for a document this category does not copy. The cursor still moves past a document that becomes null. + /// The source holds no document with the id in , which means the cursor and the database no longer belong together. + public static async IAsyncEnumerable ByPrefix( + RavenReadOnlySourceLifecycle lifecycle, + string categoryId, + string databaseName, + string prefix, + string? resumeAfter, + int batchSize, + Func, MigrationRow?> project, + [EnumeratorCancellation] CancellationToken cancellationToken = default) where TDocument : class + { + using var session = lifecycle.OpenSession(databaseName); + + // RavenDB starts a stream after whatever id it is given, even one that does not exist, so a stale cursor would skip rows nothing has copied. + if (resumeAfter is not null && !await session.Advanced.ExistsAsync(resumeAfter, cancellationToken)) + { + throw new InvalidOperationException( + $"The migration cannot resume the '{categoryId}' category after '{resumeAfter}', because the RavenDB database '{databaseName}' holds no document with that id. The cursor is saved in the target database and names a source document, so restoring RavenDB from a backup or changing which database it reads separates the two. Point the migration back at the RavenDB database this copy started from and restart."); + } + + await using var enumerator = await session.Advanced.StreamAsync( + prefix, startAfter: resumeAfter, token: cancellationToken); + + var rows = new List(batchSize); + var cursor = resumeAfter; + var lastYielded = resumeAfter; + + while (await enumerator.MoveNextAsync()) + { + // The cursor moves on every document the stream hands back, even one that becomes no row, so a resume never walks it again. + cursor = enumerator.Current.Id; + + if (project(enumerator.Current) is { } row) + { + rows.Add(row); + } + + if (rows.Count == batchSize) + { + yield return new MigrationBatch(rows, cursor); + rows = new List(batchSize); + lastYielded = cursor; + } + } + + // A tail whose documents all became no row still goes out, empty, so that the cursor past them is saved. + if (rows.Count > 0 || cursor != lastYielded) + { + yield return new MigrationBatch(rows, cursor!); + } + } + + public static MigrationRow WholeDocument(StreamResult result) where TDocument : class => + new(result.Id, result.Document, NoMetadata); + + static readonly IReadOnlyDictionary NoMetadata = ReadOnlyDictionary.Empty; +} diff --git a/src/ServiceControl.Persistence.RavenDB/DataMigration/RavenMigrationSource.cs b/src/ServiceControl.Persistence.RavenDB/DataMigration/RavenMigrationSource.cs index cc3c59a998..8cb44a7cf7 100644 --- a/src/ServiceControl.Persistence.RavenDB/DataMigration/RavenMigrationSource.cs +++ b/src/ServiceControl.Persistence.RavenDB/DataMigration/RavenMigrationSource.cs @@ -3,6 +3,7 @@ namespace ServiceControl.Persistence.RavenDB.DataMigration; using System; +using System.Collections.Frozen; using System.Collections.Generic; using System.Linq; using System.Threading; @@ -11,11 +12,25 @@ namespace ServiceControl.Persistence.RavenDB.DataMigration; using Raven.Client.Documents.Operations; using Raven.Client.ServerWide.Operations; using ServiceControl.Persistence.DataMigration; +using ServiceControl.Persistence.RavenDB.DataMigration.Readers; +/// +/// Reads a migration out of RavenDB. It holds one reader per category and knows nothing about any of them +/// beyond that, so a category this build cannot read is simply absent from +/// and never reaches the engine. +/// sealed class RavenMigrationSource(RavenReadOnlySourceLifecycle lifecycle) : IMigrationSource { + readonly FrozenDictionary readers = new IMigrationCategoryReader[] + { + new KnownEndpointsReader(lifecycle), + new EndpointSettingsReader(lifecycle) + }.ToFrozenDictionary(reader => reader.CategoryId, StringComparer.Ordinal); + public Task Open(CancellationToken cancellationToken = default) => lifecycle.Open(cancellationToken); + public IReadOnlyList ContributedChecks() => []; + public async Task Describe(CancellationToken cancellationToken = default) { var settings = lifecycle.Settings; @@ -47,18 +62,37 @@ public async Task> Inventory(Cancel return entries; } - public Task Count(MigrationCategory category, CancellationToken cancellationToken = default) => - throw new NotSupportedException($"The RavenDB migration source cannot count category {category.Id} yet"); + // Counted through the reader the copy uses, not from RavenDB's collection statistics, because a collection does not map one to one onto a category. + public async Task Count(MigrationCategory category, CancellationToken cancellationToken = default) + { + var total = 0L; + + await foreach (var batch in Read(category, resumeAfter: null, batchSize: CountBatchSize, cancellationToken)) + { + total += batch.Rows.Count; + } + + return total; + } public IAsyncEnumerable Read( MigrationCategory category, string? resumeAfter, int batchSize, CancellationToken cancellationToken = default) => - throw new NotSupportedException($"The RavenDB migration source cannot read category {category.Id} yet"); + readers.TryGetValue(category.Id, out var reader) + ? reader.Read(resumeAfter, batchSize, cancellationToken) + : throw new NotSupportedException($"The migration source cannot yet read the '{category.Id}' category."); public Task ReadBody(MigrationCategory category, string sourceId, CancellationToken cancellationToken = default) => throw new NotSupportedException($"The RavenDB migration source cannot read bodies for category {category.Id} yet"); + public IReadOnlyCollection SupportedCategoryIds => readers.Keys; + + public IReadOnlyDictionary DocumentTypes => readers.ToFrozenDictionary(pair => pair.Key, pair => pair.Value.DocumentType, StringComparer.Ordinal); + public ValueTask DisposeAsync() => lifecycle.DisposeAsync(); + + // Counting only adds up row counts, so this size changes nothing but how often the stream stops to hand one back. + const int CountBatchSize = 1024; } diff --git a/src/ServiceControl.Persistence.RavenDB/DataMigration/RavenReadOnlySourceLifecycle.cs b/src/ServiceControl.Persistence.RavenDB/DataMigration/RavenReadOnlySourceLifecycle.cs index 3b6f32ff55..901c76f6ac 100644 --- a/src/ServiceControl.Persistence.RavenDB/DataMigration/RavenReadOnlySourceLifecycle.cs +++ b/src/ServiceControl.Persistence.RavenDB/DataMigration/RavenReadOnlySourceLifecycle.cs @@ -5,6 +5,7 @@ namespace ServiceControl.Persistence.RavenDB.DataMigration; using System; using System.Diagnostics; using System.Net.Http; +using System.Security.Cryptography.X509Certificates; using System.Threading; using System.Threading.Tasks; using Microsoft.Extensions.Hosting; @@ -18,6 +19,11 @@ namespace ServiceControl.Persistence.RavenDB.DataMigration; using ServiceControl.Configuration; using ServiceControl.RavenDB; +/// +/// The connection to the RavenDB database a migration reads from, and the embedded server it starts when that +/// database is a data directory rather than a URL. Nothing here writes, so the old database stays exactly as it +/// was and the copy can still be thrown away. +/// sealed class RavenReadOnlySourceLifecycle(RavenPersisterSettings settings, SettingsRootNamespace settingsRoot) : IAsyncDisposable { public RavenPersisterSettings Settings => settings; @@ -97,7 +103,7 @@ async Task EnsureReadable(string databaseName, string settingKey, CancellationTo } catch (Exception e) when (e is not DatabaseLoadTimeoutException) { - throw new InvalidOperationException($"The RavenDB migration source at {Located()} has a database named '{databaseName}', from the '{settingKey}' setting, but could not load it.", e); + throw new InvalidOperationException($"The RavenDB migration source at {Located()} has a database named '{databaseName}', from the '{settingKey}' setting, but could not load it: {e.Message}", e); } } } @@ -110,8 +116,8 @@ internal static string Located(RavenPersisterSettings sourceSettings, SettingsRo async Task StartEmbedded(CancellationToken cancellationToken) { - // A dynamic query is a POST to /queries, which the request guard allows and which builds an auto-index - // on the customer's fallback database. This makes the server refuse it rather than trusting every reader. + // A dynamic query is a POST to /queries, which the read-only guard allows, and it builds an auto-index on the + // customer's database. Nothing on the client can tell those queries apart, so the server is told to refuse them. var configuration = new EmbeddedDatabaseConfiguration(settings.ServerUrl, settings.DatabaseName, settings.DatabasePath, settings.LogPath, settings.LogsMode) { DisableAutoIndexCreation = true }; embedded = EmbeddedDatabase.Start(configuration, lifetime); @@ -126,7 +132,7 @@ async Task StartEmbedded(CancellationToken cancellationToken) } catch (Exception e) { - throw new InvalidOperationException($"The RavenDB migration source could not start a server for the embedded database at {Located()}. A ServiceControl instance still running against that data directory is the usual cause: stop it, run the report, then start it again.", e); + throw new InvalidOperationException($"The RavenDB migration source could not start a server for the embedded database at {Located()}: {e.Message} A ServiceControl instance still holding that data directory is one cause, and stopping it lets the report run; a port already in use or a missing RavenDB server are others, which the message above tells apart.", e); } } @@ -142,6 +148,7 @@ IDocumentStore Connect(string serverUrl) if (!settings.UseEmbeddedServer) { store.Certificate = RavenClientCertificate.FindClientCertificate(settings); + RefuseUnusableCertificate(store.Certificate); } store.OnBeforeRequest += RefuseWrite; @@ -149,6 +156,30 @@ IDocumentStore Connect(string serverUrl) return store.Initialize(); } + // An unusable certificate otherwise surfaces on the first request as a refused connection, which says nothing about the cause. + void RefuseUnusableCertificate(X509Certificate2? certificate) + { + if (certificate is null) + { + if (settings.ConnectionString.StartsWith("https://", StringComparison.OrdinalIgnoreCase)) + { + throw new InvalidOperationException($"The RavenDB migration source at '{settings.ConnectionString}' is secured but no client certificate is configured. Set '{settingsRoot}/{RavenBootstrapper.ClientCertificatePathKey}' or '{settingsRoot}/{RavenBootstrapper.ClientCertificateBase64Key}'."); + } + + return; + } + + // X509Certificate2 reports both bounds in local time, so they move to UTC before the comparison. + var notBefore = certificate.NotBefore.ToUniversalTime(); + var notAfter = certificate.NotAfter.ToUniversalTime(); + var now = DateTime.UtcNow; + + if (now < notBefore || now > notAfter) + { + throw new InvalidOperationException($"The RavenDB client certificate '{certificate.Subject}' is valid from {notBefore:u} to {notAfter:u}, which does not include now."); + } + } + static void RefuseWrite(object? sender, BeforeRequestEventArgs e) { if (IsRead(e.Request.Method, new Uri(e.Url).AbsolutePath)) diff --git a/src/ServiceControl.Persistence.RavenDB/DataMigration/Readers/EndpointSettingsReader.cs b/src/ServiceControl.Persistence.RavenDB/DataMigration/Readers/EndpointSettingsReader.cs new file mode 100644 index 0000000000..087194c67c --- /dev/null +++ b/src/ServiceControl.Persistence.RavenDB/DataMigration/Readers/EndpointSettingsReader.cs @@ -0,0 +1,19 @@ +#nullable enable + +namespace ServiceControl.Persistence.RavenDB.DataMigration.Readers; + +using System.Collections.Generic; +using System.Threading; +using ServiceControl.Persistence.DataMigration; + +/// +/// Reads the per-endpoint settings, whole, out of the primary database. The collection holds one document per +/// endpoint plus one with an empty name, which carries the default for every endpoint. +/// +sealed class EndpointSettingsReader(RavenReadOnlySourceLifecycle lifecycle) : MigrationCategoryReader(lifecycle) +{ + public override string CategoryId => MigrationCategoryIds.EndpointSettings; + + public override IAsyncEnumerable Read(string? resumeAfter, int batchSize, CancellationToken cancellationToken = default) => + WholeDocuments(EndpointSettingsStore.CollectionName + "/", resumeAfter, batchSize, cancellationToken); +} diff --git a/src/ServiceControl.Persistence.RavenDB/DataMigration/Readers/KnownEndpointsReader.cs b/src/ServiceControl.Persistence.RavenDB/DataMigration/Readers/KnownEndpointsReader.cs new file mode 100644 index 0000000000..3372368c5f --- /dev/null +++ b/src/ServiceControl.Persistence.RavenDB/DataMigration/Readers/KnownEndpointsReader.cs @@ -0,0 +1,18 @@ +#nullable enable + +namespace ServiceControl.Persistence.RavenDB.DataMigration.Readers; + +using System.Collections.Generic; +using System.Threading; +using ServiceControl.Persistence.DataMigration; + +/// +/// Reads the endpoints ServiceControl has heard from, whole, out of the primary database. +/// +sealed class KnownEndpointsReader(RavenReadOnlySourceLifecycle lifecycle) : MigrationCategoryReader(lifecycle) +{ + public override string CategoryId => MigrationCategoryIds.KnownEndpoints; + + public override IAsyncEnumerable Read(string? resumeAfter, int batchSize, CancellationToken cancellationToken = default) => + WholeDocuments(RavenMonitoringDataStore.KnownEndpointsCollectionName + "/", resumeAfter, batchSize, cancellationToken); +} diff --git a/src/ServiceControl.Persistence.RavenDB/EndpointSettingsStore.cs b/src/ServiceControl.Persistence.RavenDB/EndpointSettingsStore.cs index 0792bcd8ff..5c88f09c5a 100644 --- a/src/ServiceControl.Persistence.RavenDB/EndpointSettingsStore.cs +++ b/src/ServiceControl.Persistence.RavenDB/EndpointSettingsStore.cs @@ -47,5 +47,5 @@ public async Task UpdateEndpointSettings(EndpointSettings settings, Cancellation await session.SaveChangesAsync(cancellationToken); } - const string CollectionName = "EndpointSettings"; + internal const string CollectionName = "EndpointSettings"; } \ No newline at end of file diff --git a/src/ServiceControl.Persistence.Tests.InMemory/PersistenceTestsContext.cs b/src/ServiceControl.Persistence.Tests.InMemory/PersistenceTestsContext.cs index a23b266f1b..dc2bcb5fc6 100644 --- a/src/ServiceControl.Persistence.Tests.InMemory/PersistenceTestsContext.cs +++ b/src/ServiceControl.Persistence.Tests.InMemory/PersistenceTestsContext.cs @@ -24,6 +24,8 @@ public Task Setup(IHostApplicationBuilder hostBuilder) return Task.CompletedTask; } + public Task InstallSchema(IHost host) => Task.CompletedTask; + public Task PostSetup(IHost host) => Task.CompletedTask; public Task TearDown() => Task.CompletedTask; diff --git a/src/ServiceControl.Persistence.Tests.PostgreSql/PersistenceTestsContext.cs b/src/ServiceControl.Persistence.Tests.PostgreSql/PersistenceTestsContext.cs index 4ce91496d5..b99a642473 100644 --- a/src/ServiceControl.Persistence.Tests.PostgreSql/PersistenceTestsContext.cs +++ b/src/ServiceControl.Persistence.Tests.PostgreSql/PersistenceTestsContext.cs @@ -51,7 +51,7 @@ public async Task Setup(IHostApplicationBuilder hostBuilder) hostBuilder.Services.AddSingleton(FakeTime); } - public async Task PostSetup(IHost host) + public async Task InstallSchema(IHost host) { this.host = host; @@ -59,6 +59,8 @@ public async Task PostSetup(IHost host) await scope.ServiceProvider.GetRequiredService().ApplyMigrations(); } + public Task PostSetup(IHost host) => Task.CompletedTask; + public async Task TearDown() { DeleteBodyStorage(); diff --git a/src/ServiceControl.Persistence.Tests.RavenDB/DataMigration/KnownEndpointSourceTests.cs b/src/ServiceControl.Persistence.Tests.RavenDB/DataMigration/KnownEndpointSourceTests.cs new file mode 100644 index 0000000000..66b7214ec9 --- /dev/null +++ b/src/ServiceControl.Persistence.Tests.RavenDB/DataMigration/KnownEndpointSourceTests.cs @@ -0,0 +1,66 @@ +namespace ServiceControl.Persistence.Tests.RavenDB.DataMigration; + +using System; +using System.Linq; +using System.Threading.Tasks; +using NUnit.Framework; +using ServiceControl.Operations; +using ServiceControl.Persistence; +using ServiceControl.Persistence.DataMigration; + +class KnownEndpointSourceTests : RavenMigrationSourceTestBase +{ + [Test] + public async Task Reads_known_endpoints_by_the_plural_prefix_in_batches_with_their_monitored_flag() + { + using (var session = await SessionProvider.OpenSession()) + { + foreach (var (name, monitored) in new[] { ("Sales.Orders", true), ("Billing", false), ("Shipping", false) }) + { + var endpoint = new KnownEndpoint + { + EndpointDetails = new EndpointDetails { Name = name, HostId = Guid.NewGuid(), Host = "HOST01" }, + HostDisplayName = "HOST01", + Monitored = monitored + }; + + await session.StoreAsync(endpoint, $"KnownEndpoints/{endpoint.EndpointDetails.GetDeterministicId()}"); + } + + await session.SaveChangesAsync(); + } + + await using var source = await OpenMigrationSource(); + var category = MigrationCategoryRegistry.All.Single(entry => entry.Id == MigrationCategoryIds.KnownEndpoints); + + var batches = await CollectBatches(source, category, batchSize: 2); + + Assert.That(batches.Select(batch => batch.Rows.Count), Is.EqualTo(new[] { 2, 1 })); + Assert.That(batches.SelectMany(batch => batch.Rows).Count(row => ((KnownEndpoint)row.Document).Monitored), Is.EqualTo(1)); + } + + // KnownEndpointsWriter skips these as RequiredValueMissing, which it can only do if the reader hands them over despite the required members. + [Test] + public async Task Reads_a_known_endpoint_missing_a_required_value_with_that_value_null() + { + using (var session = await SessionProvider.OpenSession()) + { + await session.StoreAsync(new { EndpointDetails = new { HostId = Guid.NewGuid(), Host = "HOST01" }, Monitored = false }, "KnownEndpoints/no-name"); + await session.StoreAsync(new { EndpointDetails = new { Name = "Sales.Orders", HostId = Guid.NewGuid() }, Monitored = false }, "KnownEndpoints/no-host"); + await session.StoreAsync(new { Monitored = false }, "KnownEndpoints/no-details"); + await session.SaveChangesAsync(); + } + + await using var source = await OpenMigrationSource(); + var category = MigrationCategoryRegistry.All.Single(entry => entry.Id == MigrationCategoryIds.KnownEndpoints); + + var endpoints = (await CollectBatches(source, category)).SelectMany(batch => batch.Rows).ToDictionary(row => row.SourceId, row => (KnownEndpoint)row.Document); + + using (Assert.EnterMultipleScope()) + { + Assert.That(endpoints["KnownEndpoints/no-name"].EndpointDetails.Name, Is.Null); + Assert.That(endpoints["KnownEndpoints/no-host"].EndpointDetails.Host, Is.Null); + Assert.That(endpoints["KnownEndpoints/no-details"].EndpointDetails, Is.Null); + } + } +} diff --git a/src/ServiceControl.Persistence.Tests.RavenDB/DataMigration/MigrationCategoryReadersTests.cs b/src/ServiceControl.Persistence.Tests.RavenDB/DataMigration/MigrationCategoryReadersTests.cs new file mode 100644 index 0000000000..9fb6081bb3 --- /dev/null +++ b/src/ServiceControl.Persistence.Tests.RavenDB/DataMigration/MigrationCategoryReadersTests.cs @@ -0,0 +1,37 @@ +namespace ServiceControl.Persistence.Tests.RavenDB.DataMigration; + +using System; +using System.Linq; +using System.Threading.Tasks; +using NUnit.Framework; +using ServiceControl.Persistence.DataMigration; +using ServiceControl.Persistence.RavenDB.DataMigration; + +[TestFixture] +class MigrationCategoryReadersTests : RavenMigrationSourceTestBase +{ + [Test] + public async Task Every_reader_claims_a_category_the_registry_knows() + { + await using var source = (RavenMigrationSource)await OpenMigrationSource(); + + var unknown = source.SupportedCategoryIds.Where(id => MigrationCategoryRegistry.Find(id) is null).ToArray(); + + Assert.That(unknown, Is.Empty, "a reader for an id no category has is dead code the engine can never reach"); + } + + [Test] + public void Every_reader_derives_from_the_typed_base() + { + var readers = typeof(IMigrationCategoryReader).Assembly.GetTypes() + .Where(type => type is { IsClass: true, IsAbstract: false } && typeof(IMigrationCategoryReader).IsAssignableFrom(type)) + .ToArray(); + + Assert.That(readers, Is.Not.Empty, "the test proves nothing if it finds no reader"); + Assert.That(readers.Where(type => !DerivesFrom(type, typeof(MigrationCategoryReader<>))).Select(type => type.Name), Is.Empty, + "a reader that implements the interface directly can declare a DocumentType it does not yield, which the pairing test cannot see"); + } + + static bool DerivesFrom(Type type, Type genericBase) => + type.BaseType is { } baseType && ((baseType.IsGenericType && baseType.GetGenericTypeDefinition() == genericBase) || DerivesFrom(baseType, genericBase)); +} diff --git a/src/ServiceControl.Persistence.Tests.RavenDB/DataMigration/RavenMigrationSourceTestBase.cs b/src/ServiceControl.Persistence.Tests.RavenDB/DataMigration/RavenMigrationSourceTestBase.cs new file mode 100644 index 0000000000..358c3f9943 --- /dev/null +++ b/src/ServiceControl.Persistence.Tests.RavenDB/DataMigration/RavenMigrationSourceTestBase.cs @@ -0,0 +1,67 @@ +namespace ServiceControl.Persistence.Tests.RavenDB.DataMigration; + +using System; +using System.Collections.Generic; +using System.Threading.Tasks; +using NUnit.Framework; +using ServiceControl.Configuration; +using ServiceControl.Persistence.DataMigration; +using ServiceControl.Persistence.RavenDB; + +abstract class RavenMigrationSourceTestBase : RavenPersistenceTestBase +{ + static readonly SettingsRootNamespace SettingsRoot = new("ServiceControl"); + static readonly object SettingsGate = new(); + + protected async Task OpenMigrationSource() + { + var settings = (RavenPersisterSettings)PersistenceSettings; + (string Name, string Value)[] variables = + [ + ("SERVICECONTROL_RAVENDB_CONNECTIONSTRING", settings.ConnectionString), + ("SERVICECONTROL_RAVENDB_DATABASENAME", settings.DatabaseName), + ("SERVICECONTROL_ERRORRETENTIONPERIOD", settings.ErrorRetentionPeriod.ToString()), + ("LICENSINGCOMPONENT_RAVENDB_THROUGHPUTDATABASENAME", settings.ThroughputDatabaseName) + ]; + + IMigrationSource source; + + // Environment variables are process wide and these tests run in parallel, so without the gate a source + // reads whichever test set them last and opens another test's database. + lock (SettingsGate) + { + try + { + foreach (var (name, value) in variables) + { + Environment.SetEnvironmentVariable(name, value); + } + + source = new RavenPersistenceConfiguration().CreateSource(SettingsRoot); + } + finally + { + foreach (var (name, _) in variables) + { + Environment.SetEnvironmentVariable(name, null); + } + } + } + + await source.Open(); + + return source; + } + + protected static async Task> CollectBatches(IMigrationSource source, MigrationCategory category, string resumeAfter = null, int batchSize = 100) + { + var batches = new List(); + + await foreach (var batch in source.Read(category, resumeAfter, batchSize, TestContext.CurrentContext.CancellationToken)) + { + batches.Add(batch); + } + + return batches; + } +} diff --git a/src/ServiceControl.Persistence.Tests.RavenDB/DataMigration/RavenMigrationSourceTests.cs b/src/ServiceControl.Persistence.Tests.RavenDB/DataMigration/RavenMigrationSourceTests.cs new file mode 100644 index 0000000000..c34641d535 --- /dev/null +++ b/src/ServiceControl.Persistence.Tests.RavenDB/DataMigration/RavenMigrationSourceTests.cs @@ -0,0 +1,145 @@ +namespace ServiceControl.Persistence.Tests.RavenDB.DataMigration; + +using System; +using System.Collections.Generic; +using System.Linq; +using System.Threading.Tasks; +using NUnit.Framework; +using Raven.Client.Documents.Operations.Indexes; +using ServiceControl.Operations; +using ServiceControl.Persistence.DataMigration; + +class RavenMigrationSourceTests : RavenMigrationSourceTestBase +{ + [Test] + public async Task Reads_endpoint_settings_in_document_id_order_in_batches_of_the_requested_size() + { + await SeedEndpointSettings(5); + await using var source = await OpenMigrationSource(); + + var batches = await CollectBatches(source, EndpointSettingsCategory, batchSize: 2); + + Assert.Multiple(() => + { + Assert.That(batches.Select(batch => batch.Rows.Count), Is.EqualTo(new[] { 2, 2, 1 }), "A batch never exceeds the requested size, and the trailing partial batch is still yielded."); + Assert.That(batches.SelectMany(batch => batch.Rows).Select(row => row.SourceId), Is.Ordered); + }); + } + + [Test] + public async Task Every_row_is_identified_by_the_document_id_the_persister_wrote_it_under() + { + await SeedEndpointSettings(5); + await using var source = await OpenMigrationSource(); + + var rows = (await CollectBatches(source, EndpointSettingsCategory)).SelectMany(batch => batch.Rows).ToList(); + + Assert.Multiple(() => + { + Assert.That(rows, Has.Count.EqualTo(5)); + Assert.That(rows.Select(row => row.SourceId), Has.All.StartWith("EndpointSettings/"), "The cursor the engine checkpoints is a source document id, so a row identified by anything else cannot be resumed after."); + }); + } + + [Test] + public async Task Resuming_after_a_cursor_yields_only_what_follows_it() + { + await SeedEndpointSettings(5); + await using var source = await OpenMigrationSource(); + + var firstBatch = (await CollectBatches(source, EndpointSettingsCategory, batchSize: 2))[0]; + var resumed = (await CollectBatches(source, EndpointSettingsCategory, resumeAfter: firstBatch.Cursor, batchSize: 2)).SelectMany(batch => batch.Rows).ToList(); + + Assert.That(resumed, Has.Count.EqualTo(3)); + Assert.That(resumed.Select(row => row.SourceId), Has.No.Member(firstBatch.Rows[0].SourceId)); + Assert.That(resumed.Select(row => row.SourceId), Is.Ordered); + } + + [Test] + public async Task Resuming_after_a_cursor_this_source_never_issued_refuses_instead_of_starting_past_it() + { + // The cursor is saved in the target database, so a restored backup or a corrected database name leaves + // one naming a document this source does not have. + await SeedEndpointSettings(3); + await using var source = await OpenMigrationSource(); + + var refusal = Assert.ThrowsAsync(() => CollectBatches(source, EndpointSettingsCategory, resumeAfter: "EndpointSettings/zzz-from-another-database")); + + Assert.Multiple(() => + { + Assert.That(refusal.Message, Does.Contain(EndpointSettingsCategory.Id)); + Assert.That(refusal.Message, Does.Contain("EndpointSettings/zzz-from-another-database")); + Assert.That(refusal.Message, Does.Contain(DatabaseName), "the operator has to be told which database the cursor was looked for in"); + }); + } + + [Test] + public async Task Counts_every_row_the_category_would_read() + { + await SeedEndpointSettings(5); + await using var source = await OpenMigrationSource(); + + Assert.That(await source.Count(EndpointSettingsCategory), Is.EqualTo(5)); + } + + [Test] + public async Task Counts_a_category_with_nothing_in_it_as_zero() + { + await using var source = await OpenMigrationSource(); + + Assert.That(await source.Count(EndpointSettingsCategory), Is.Zero, "A category with nothing to copy has to be told apart from one this source declines to count."); + } + + [Test] + public async Task Reading_every_category_creates_no_index() + { + await using var source = await OpenMigrationSource(); + + Assert.That(source.SupportedCategoryIds, Is.SubsetOf(Seeds.Keys), "A category read with nothing in it cannot show whether its reader builds an index, so a new reader needs its seed adding here."); + + foreach (var categoryId in source.SupportedCategoryIds) + { + await Seeds[categoryId](this); + } + + var before = await IndexNames(); + + foreach (var categoryId in source.SupportedCategoryIds) + { + await CollectBatches(source, MigrationCategoryRegistry.Find(categoryId)); + } + + var after = await IndexNames(); + + Assert.Multiple(() => + { + Assert.That(after, Is.EquivalentTo(before), "A reader that queries without naming a static index has RavenDB build one and index the whole collection on the customer's live database."); + Assert.That(after, Has.None.StartWith("Auto/")); + }); + } + + // Keyed by category id so a reader added without seed data fails the index test by name instead of passing + // over an empty collection. + static readonly Dictionary> Seeds = new() + { + [MigrationCategoryIds.EndpointSettings] = tests => tests.SeedEndpointSettings(2), + [MigrationCategoryIds.KnownEndpoints] = tests => tests.MonitoringDataStore.CreateIfNotExists(new EndpointDetails { Name = "Sales.Orders", HostId = Guid.NewGuid(), Host = "HOST01" }) + }; + + string DatabaseName => ((RavenPersisterSettings)PersistenceSettings).DatabaseName; + + async Task IndexNames() => + await DocumentStore.Maintenance.SendAsync(new GetIndexNamesOperation(0, IndexNamePageSize), TestContext.CurrentContext.CancellationToken); + + const int IndexNamePageSize = 1024; + + static readonly MigrationCategory EndpointSettingsCategory = MigrationCategoryRegistry.Find("EndpointSettings"); + + async Task SeedEndpointSettings(int count) + { + for (var index = 0; index < count; index++) + { + await EndpointSettingsStore.UpdateEndpointSettings(new EndpointSettings { Name = $"Endpoint{index}", TrackInstances = true }); + } + } +} diff --git a/src/ServiceControl.Persistence.Tests.RavenDB/DataMigration/ReadOnlySourceLifecycleTests.cs b/src/ServiceControl.Persistence.Tests.RavenDB/DataMigration/ReadOnlySourceLifecycleTests.cs index 859e02c505..ba3f892ffe 100644 --- a/src/ServiceControl.Persistence.Tests.RavenDB/DataMigration/ReadOnlySourceLifecycleTests.cs +++ b/src/ServiceControl.Persistence.Tests.RavenDB/DataMigration/ReadOnlySourceLifecycleTests.cs @@ -88,6 +88,20 @@ public async Task Opening_the_source_creates_no_database() Assert.That(await bootstrapStore.Maintenance.Server.SendAsync(new GetDatabaseRecordOperation(absentThroughput)), Is.Null, "Opening a migration source must not create a database that was missing."); } + [Test] + public async Task Opening_the_source_refuses_a_missing_primary_database() + { + var absentPrimary = $"{databaseName}-absent"; + sourceSettings.DatabaseName = absentPrimary; + + await using var lifecycle = new RavenReadOnlySourceLifecycle(sourceSettings, SettingsRoot); + + var exception = Assert.ThrowsAsync(async () => await lifecycle.Open()); + + Assert.That(exception.Message, Does.Contain(absentPrimary).And.Contain("ServiceControl/RavenDB/DatabaseName")); + Assert.That(await bootstrapStore.Maintenance.Server.SendAsync(new GetDatabaseRecordOperation(absentPrimary)), Is.Null, "Opening a migration source must not create a primary database that was missing."); + } + [Test] public async Task Opening_the_source_writes_no_database_settings() { diff --git a/src/ServiceControl.Persistence.Tests.RavenDB/PersistenceTestsContext.cs b/src/ServiceControl.Persistence.Tests.RavenDB/PersistenceTestsContext.cs index 5bcae71f20..aca5f06236 100644 --- a/src/ServiceControl.Persistence.Tests.RavenDB/PersistenceTestsContext.cs +++ b/src/ServiceControl.Persistence.Tests.RavenDB/PersistenceTestsContext.cs @@ -50,6 +50,8 @@ public async Task Setup(IHostApplicationBuilder hostBuilder) persistence.AddInstaller(hostBuilder.Services); } + public Task InstallSchema(IHost host) => Task.CompletedTask; + public async Task PostSetup(IHost host) { DocumentStore = await host.Services.GetRequiredService().GetDocumentStore(); diff --git a/src/ServiceControl.Persistence.Tests.SqlServer/EndpointSettingsKeyCollationTests.cs b/src/ServiceControl.Persistence.Tests.SqlServer/EndpointSettingsKeyCollationTests.cs new file mode 100644 index 0000000000..8a90e4965c --- /dev/null +++ b/src/ServiceControl.Persistence.Tests.SqlServer/EndpointSettingsKeyCollationTests.cs @@ -0,0 +1,172 @@ +namespace ServiceControl.Persistence.Tests; + +using System; +using System.Collections.Generic; +using System.Linq; +using System.Threading.Tasks; +using Microsoft.EntityFrameworkCore; +using Microsoft.EntityFrameworkCore.Infrastructure; +using Microsoft.EntityFrameworkCore.Storage; +using Microsoft.Extensions.DependencyInjection; +using Microsoft.Extensions.Logging; +using NUnit.Framework; +using ServiceControl.Operations; +using ServiceControl.Persistence.DataMigration; +using ServiceControl.Persistence.EFCore.DataMigration; +using ServiceControl.Persistence.EFCore.DbContexts; +using ServiceControl.Persistence.EFCore.Entities; +using ServiceControl.Persistence.EFCore.Infrastructure; + +class EndpointSettingsKeyCollationTests : PersistenceTestBase +{ + readonly TargetWarnings targetWarnings = new(); + + public EndpointSettingsKeyCollationTests() => RegisterServices = services => services.AddSingleton(targetWarnings); + + [Test] + public async Task A_case_merge_inside_one_batch_logs_both_names_and_the_one_kept() + { + await OpenACaseInsensitiveTarget("Sales", "sales"); + + await Write(("EndpointSettings/1", "Sales"), ("EndpointSettings/2", "sales")); + + Assert.That(targetWarnings.Messages.Count(message => message.Contains("'sales'") && message.Contains("kept 'Sales'")), Is.EqualTo(1), string.Join(Environment.NewLine, targetWarnings.Messages)); + } + + [Test] + public async Task A_case_merge_with_a_row_an_earlier_batch_wrote_is_logged_and_a_row_written_again_is_not() + { + await OpenACaseInsensitiveTarget("Sales", "sales"); + + await Write(("EndpointSettings/1", "Sales")); + await Write(("EndpointSettings/1", "Sales"), ("EndpointSettings/2", "sales")); + + Assert.That(targetWarnings.Messages, Has.Exactly(1).Contains("kept 'Sales'"), string.Join(Environment.NewLine, targetWarnings.Messages)); + } + + async Task OpenACaseInsensitiveTarget(params string[] knownEndpointNames) + { + using (var scope = ServiceProvider.CreateScope()) + { + var dbContext = scope.ServiceProvider.GetRequiredService(); + var entityType = dbContext.Model.FindEntityType(typeof(EndpointSettingsEntity)); + var table = dbContext.GetService().DelimitIdentifier(entityType.GetTableName(), entityType.GetSchema()); + + // Set on the column, because the merge follows the column's collation and the test server's default can be either. + var ignoreCase = $""" + ALTER TABLE {table} DROP CONSTRAINT [PK_EndpointSettings]; + ALTER TABLE {table} ALTER COLUMN [Name] nvarchar(450) COLLATE Latin1_General_CI_AS NOT NULL; + ALTER TABLE {table} ADD CONSTRAINT [PK_EndpointSettings] PRIMARY KEY ([Name]); + """; + + await dbContext.Database.ExecuteSqlRawAsync(ignoreCase); + } + + foreach (var name in knownEndpointNames) + { + await MonitoringDataStore.CreateIfNotExists(new EndpointDetails { Name = name, HostId = Guid.NewGuid(), Host = "HOST01" }); + } + + await ServiceProvider.GetRequiredService().Open(); + } + + async Task Write(params (string SourceId, string Name)[] rows) + { + var category = MigrationCategoryRegistry.All.Single(entry => entry.Id == MigrationCategoryIds.EndpointSettings); + var batch = new MigrationBatch( + [.. rows.Select(row => new MigrationRow(row.SourceId, new EndpointSettings { Name = row.Name, TrackInstances = true }, new Dictionary()))], + rows[^1].SourceId); + var checkpoint = await ServiceProvider.GetRequiredService().Read(category.Id) + ?? new MigrationCheckpoint(category.Id, MigrationCategoryState.InProgress, null, 0, 0, null, null, null, null, null, null); + + await ServiceProvider.GetRequiredService().Write(category, batch, checkpoint); + } + + [Test] + public async Task The_key_columns_own_collation_decides_a_merge_when_the_database_default_disagrees() + { + bool databaseIgnoresCase; + + using (var scope = ServiceProvider.CreateScope()) + { + var dbContext = scope.ServiceProvider.GetRequiredService(); + var database = dbContext.Database; + + databaseIgnoresCase = await database + .SqlQuery($"SELECT CONVERT(int, DATABASEPROPERTYEX(DB_NAME(), 'ComparisonStyle')) & 1 AS [Value]") + .SingleAsync() == 1; + + // Named from the model, because the table sits in the configured schema when the persister has one. + var entityType = dbContext.Model.FindEntityType(typeof(EndpointSettingsEntity)); + var table = dbContext.GetService().DelimitIdentifier(entityType.GetTableName(), entityType.GetSchema()); + + // The column is given the opposite of the database default, because a test where the two agree cannot show which one decided. + var recollate = $""" + DECLARE @columnCollation sysname = IIF(CONVERT(int, DATABASEPROPERTYEX(DB_NAME(), 'ComparisonStyle')) & 1 = 1, N'Latin1_General_CS_AS', N'Latin1_General_CI_AS'); + ALTER TABLE {table} DROP CONSTRAINT [PK_EndpointSettings]; + EXEC (N'ALTER TABLE {table} ALTER COLUMN [Name] nvarchar(450) COLLATE ' + @columnCollation + N' NOT NULL'); + ALTER TABLE {table} ADD CONSTRAINT [PK_EndpointSettings] PRIMARY KEY ([Name]); + """; + + await database.ExecuteSqlRawAsync(recollate); + } + + foreach (var name in new[] { "Sales", "sales" }) + { + await MonitoringDataStore.CreateIfNotExists(new EndpointDetails { Name = name, HostId = Guid.NewGuid(), Host = "HOST01" }); + } + + var target = ServiceProvider.GetRequiredService(); + await target.Open(); + + var category = MigrationCategoryRegistry.All.Single(entry => entry.Id == MigrationCategoryIds.EndpointSettings); + var batch = new MigrationBatch( + [ + new MigrationRow("EndpointSettings/1", new EndpointSettings { Name = "Sales", TrackInstances = true }, new Dictionary()), + new MigrationRow("EndpointSettings/2", new EndpointSettings { Name = "sales", TrackInstances = false }, new Dictionary()) + ], + "EndpointSettings/2"); + var checkpointToExtend = new MigrationCheckpoint(category.Id, MigrationCategoryState.InProgress, batch.Cursor, 0, 0, null, null, null, null, null, null); + + var result = await target.Write(category, batch, checkpointToExtend); + + using (Assert.EnterMultipleScope()) + { + Assert.That(result.Copied, Is.EqualTo(databaseIgnoresCase ? 2 : 1), "the column's own collation decides: two keys where it respects case, one where it ignores case, whatever the database default says"); + Assert.That( + ServiceProvider.GetRequiredService().KeyComparer(typeof(EndpointSettingsEntity), nameof(EndpointSettingsEntity.Name)).Equals("Sales", "sales"), + Is.EqualTo(!databaseIgnoresCase), + "the dry run predicts merges with this comparer, so it follows the same column the statement does"); + } + } + + // RecordingLoggerProvider lives in ServiceControl.Infrastructure.Tests, which this project does not reference. + sealed class TargetWarnings : ILoggerProvider + { + readonly System.Collections.Concurrent.ConcurrentQueue messages = new(); + + public IReadOnlyList Messages => [.. messages]; + + public ILogger CreateLogger(string categoryName) => + categoryName == typeof(EFCoreMigrationTarget).FullName ? new WarningLogger(messages) : Microsoft.Extensions.Logging.Abstractions.NullLogger.Instance; + + public void Dispose() + { + } + + sealed class WarningLogger(System.Collections.Concurrent.ConcurrentQueue messages) : ILogger + { + public IDisposable BeginScope(TState state) where TState : notnull => null; + + public bool IsEnabled(LogLevel logLevel) => logLevel == LogLevel.Warning; + + public void Log(LogLevel logLevel, EventId eventId, TState state, Exception exception, Func formatter) + { + if (logLevel == LogLevel.Warning) + { + messages.Enqueue(formatter(state, exception)); + } + } + } + } +} diff --git a/src/ServiceControl.Persistence.Tests.SqlServer/PersistenceTestsContext.cs b/src/ServiceControl.Persistence.Tests.SqlServer/PersistenceTestsContext.cs index 9ded3182bc..2dab62abe4 100644 --- a/src/ServiceControl.Persistence.Tests.SqlServer/PersistenceTestsContext.cs +++ b/src/ServiceControl.Persistence.Tests.SqlServer/PersistenceTestsContext.cs @@ -50,7 +50,7 @@ public async Task Setup(IHostApplicationBuilder hostBuilder) hostBuilder.Services.AddSingleton(FakeTime); } - public async Task PostSetup(IHost host) + public async Task InstallSchema(IHost host) { this.host = host; @@ -58,6 +58,8 @@ public async Task PostSetup(IHost host) await scope.ServiceProvider.GetRequiredService().ApplyMigrations(); } + public Task PostSetup(IHost host) => Task.CompletedTask; + public async Task TearDown() { DeleteBodyStorage(); diff --git a/src/ServiceControl.Persistence.Tests/EFCore/EFMigrationCheckpointStoreTests.cs b/src/ServiceControl.Persistence.Tests/EFCore/EFMigrationCheckpointStoreTests.cs index b96684bb33..67d7c9c307 100644 --- a/src/ServiceControl.Persistence.Tests/EFCore/EFMigrationCheckpointStoreTests.cs +++ b/src/ServiceControl.Persistence.Tests/EFCore/EFMigrationCheckpointStoreTests.cs @@ -23,7 +23,8 @@ public async Task Upsert_then_Read_round_trips_every_field() { var checkpoint = new MigrationCheckpoint( "EndpointSettings", MigrationCategoryState.CompleteWithErrors, "cursor-1", 5, 2, 40, - new Dictionary { [MigrationSkipReason.BodyUnreadable] = 2 }, Now, Now, Now, "body storage unavailable", AlreadyPresentCount: 33); + new Dictionary { [MigrationSkipReason.BodyUnreadable] = 2 }, Now, Now, Now, "body storage unavailable", AlreadyPresentCount: 33, + StartedWindowSeconds: (long)TimeSpan.FromDays(14).TotalSeconds); var saved = await Store.Upsert(checkpoint); var stored = await Store.Read("EndpointSettings"); diff --git a/src/ServiceControl.Persistence.Tests/EFCore/Migration/EndpointSettingsMigrationTargetTests.cs b/src/ServiceControl.Persistence.Tests/EFCore/Migration/EndpointSettingsMigrationTargetTests.cs new file mode 100644 index 0000000000..2826d26011 --- /dev/null +++ b/src/ServiceControl.Persistence.Tests/EFCore/Migration/EndpointSettingsMigrationTargetTests.cs @@ -0,0 +1,143 @@ +namespace ServiceControl.Persistence.Tests; + +using System.Collections.Generic; +using System.Linq; +using System.Threading.Tasks; +using Microsoft.EntityFrameworkCore; +using Microsoft.Extensions.DependencyInjection; +using NUnit.Framework; +using ServiceControl.Persistence.DataMigration; +using ServiceControl.Persistence.EFCore.DbContexts; + +class EndpointSettingsMigrationTargetTests : PersistenceTestBase +{ + [SetUp] + public Task OpenTarget() => Target.Open(); + + IMigrationTarget Target => ServiceProvider.GetRequiredService(); + + IMigrationCheckpointStore CheckpointStore => ServiceProvider.GetRequiredService(); + + static readonly MigrationCategory EndpointSettingsCategory = MigrationCategoryRegistry.All.Single(category => category.Id == MigrationCategoryIds.EndpointSettings); + + [Test] + public async Task Every_setting_in_a_batch_arrives_with_its_track_instances_value() + { + var result = await Target.Write( + EndpointSettingsCategory, + BatchOf( + ("EndpointSettings/1", new EndpointSettings { Name = "Sales", TrackInstances = true }), + ("EndpointSettings/2", new EndpointSettings { Name = "Billing", TrackInstances = false }), + ("EndpointSettings/3", new EndpointSettings { Name = "Shipping", TrackInstances = true })), + CheckpointAfter("EndpointSettings/3")); + + using (Assert.EnterMultipleScope()) + { + Assert.That(result.Copied, Is.EqualTo(3), "Copied is what the insert added, not the size of the batch"); + Assert.That(await TrackInstancesFor("Sales"), Is.True); + Assert.That(await TrackInstancesFor("Billing"), Is.False); + Assert.That(await TrackInstancesFor("Shipping"), Is.True); + } + } + + [Test] + public async Task A_name_already_in_the_target_is_left_alone_and_counted_as_already_present() + { + await EndpointSettingsStore.UpdateEndpointSettings(new EndpointSettings { Name = "Sales", TrackInstances = true }); + + var result = await Target.Write( + EndpointSettingsCategory, + BatchOf(("EndpointSettings/7", new EndpointSettings { Name = "Sales", TrackInstances = false })), + CheckpointAfter("EndpointSettings/7")); + + using (Assert.EnterMultipleScope()) + { + Assert.That(result.Copied, Is.Zero); + Assert.That(result.AlreadyPresent, Is.EqualTo(1)); + Assert.That(result.Skipped, Is.Zero, "a row the target already holds lost nothing, so it is not a skip"); + Assert.That(await TrackInstancesFor("Sales"), Is.True); + } + } + + [Test] + public async Task A_setting_whose_endpoint_is_not_known_is_copied_and_left_to_the_heartbeat_sync() + { + var result = await Target.Write( + EndpointSettingsCategory, + BatchOf( + ("EndpointSettings/1", new EndpointSettings { Name = string.Empty, TrackInstances = true }), + ("EndpointSettings/2", new EndpointSettings { Name = "Retired", TrackInstances = false })), + CheckpointAfter("EndpointSettings/2")); + + using (Assert.EnterMultipleScope()) + { + Assert.That(result.Copied, Is.EqualTo(2)); + Assert.That(result.Skipped, Is.Zero, "whether a setting outlives its endpoint is the heartbeat sync's rule, not the migration's"); + Assert.That(await TrackInstancesFor(string.Empty), Is.True); + Assert.That(await TrackInstancesFor("Retired"), Is.False); + } + } + + [Test] + public async Task A_batch_that_copies_nothing_still_saves_the_cursor() + { + await EndpointSettingsStore.UpdateEndpointSettings(new EndpointSettings { Name = "Sales", TrackInstances = true }); + await EndpointSettingsStore.UpdateEndpointSettings(new EndpointSettings { Name = "Billing", TrackInstances = true }); + + var result = await Target.Write( + EndpointSettingsCategory, + BatchOf( + ("EndpointSettings/1", new EndpointSettings { Name = "Sales", TrackInstances = false }), + ("EndpointSettings/2", new EndpointSettings { Name = "Billing", TrackInstances = false })), + CheckpointAfter("EndpointSettings/2")); + + var stored = await CheckpointStore.Read(MigrationCategoryIds.EndpointSettings); + + using (Assert.EnterMultipleScope()) + { + Assert.That(result.Copied, Is.Zero); + Assert.That(result.AlreadyPresent, Is.EqualTo(2)); + Assert.That(stored.Cursor, Is.EqualTo("EndpointSettings/2"), "a batch that copied nothing still has to commit the cursor past it, or the restart reads the same rows forever"); + } + } + + [Test] + public async Task Every_mapped_column_is_set_from_a_fully_populated_document() + { + await Target.Write(EndpointSettingsCategory, BatchOf(("EndpointSettings/1", new EndpointSettings { Name = "Sales", TrackInstances = true })), CheckpointAfter("EndpointSettings/1")); + + using var scope = ServiceProvider.CreateScope(); + var dbContext = scope.ServiceProvider.GetRequiredService(); + + MigrationEntityCoverage.AssertEveryMappedPropertyIsSet(dbContext.Model, await dbContext.EndpointSettings.AsNoTracking().SingleAsync()); + } + + static MigrationBatch BatchOf(params (string Id, object Document)[] rows) => + new([.. rows.Select(row => new MigrationRow(row.Id, row.Document, new Dictionary()))], rows[^1].Id); + + static MigrationCheckpoint CheckpointAfter(string cursor) => + new(MigrationCategoryIds.EndpointSettings, MigrationCategoryState.InProgress, cursor, 0, 0, null, null, null, null, null, null); + + async Task TrackInstancesFor(string name) => + (await EndpointSettingsStore.GetAllEndpointSettings().ToListAsync()).Single(settings => settings.Name == name).TrackInstances; + + [Test] + public async Task Two_names_in_one_batch_differing_only_in_case_are_each_copied_or_already_present() + { + var result = await Target.Write( + EndpointSettingsCategory, + BatchOf( + ("EndpointSettings/1", new EndpointSettings { Name = "Sales", TrackInstances = true }), + ("EndpointSettings/2", new EndpointSettings { Name = "sales", TrackInstances = false })), + CheckpointAfter("EndpointSettings/2")); + + using (Assert.EnterMultipleScope()) + { + Assert.That(result.Copied + result.AlreadyPresent, Is.EqualTo(2), "SQL Server's default collation makes these one key and PostgreSQL makes them two; either way neither row may throw or go uncounted"); + Assert.That(result.Skipped, Is.Zero, "a merge onto a row the batch writes loses nothing the target lacks, and a skip would count toward the halt threshold"); + Assert.That(await Target.Count(EndpointSettingsCategory), Is.EqualTo(result.Copied), "every row reported as copied is a row in the table"); + Assert.That(await TrackInstancesFor("Sales"), Is.True, "the first row in document-id order wins where the key makes the two one"); + } + } + +} diff --git a/src/ServiceControl.Persistence.Tests/EFCore/Migration/KnownEndpointsMigrationTargetTests.cs b/src/ServiceControl.Persistence.Tests/EFCore/Migration/KnownEndpointsMigrationTargetTests.cs new file mode 100644 index 0000000000..41cf3f5bd5 --- /dev/null +++ b/src/ServiceControl.Persistence.Tests/EFCore/Migration/KnownEndpointsMigrationTargetTests.cs @@ -0,0 +1,199 @@ +namespace ServiceControl.Persistence.Tests; + +using System; +using System.Collections.Generic; +using System.Linq; +using System.Threading.Tasks; +using Microsoft.EntityFrameworkCore; +using Microsoft.Extensions.DependencyInjection; +using NUnit.Framework; +using ServiceControl.Operations; +using ServiceControl.Persistence.DataMigration; +using ServiceControl.Persistence.EFCore.DataMigration; +using ServiceControl.Persistence.EFCore.DbContexts; + +class KnownEndpointsMigrationTargetTests : PersistenceTestBase +{ + [SetUp] + public Task OpenTarget() => Target.Open(); + + IMigrationTarget Target => ServiceProvider.GetRequiredService(); + + static readonly MigrationCategory KnownEndpointsCategory = MigrationCategoryRegistry.All.Single(category => category.Id == MigrationCategoryIds.KnownEndpoints); + + [Test] + public async Task A_known_endpoint_is_written_with_its_monitored_flag() + { + var endpoint = new KnownEndpoint + { + EndpointDetails = new EndpointDetails { Name = "Sales.Orders", HostId = Guid.NewGuid(), Host = "SALES01" }, + HostDisplayName = "SALES01", + Monitored = true + }; + var sourceId = $"KnownEndpoints/{endpoint.EndpointDetails.GetDeterministicId()}"; + + var result = await Target.Write(KnownEndpointsCategory, BatchOf((sourceId, endpoint)), CheckpointAfter(sourceId)); + + var stored = (await MonitoringDataStore.GetAllKnownEndpoints()).Single(); + + using (Assert.EnterMultipleScope()) + { + Assert.That(result.Copied, Is.EqualTo(1)); + Assert.That(stored.Monitored, Is.True); + Assert.That(stored.EndpointDetails.Name, Is.EqualTo("Sales.Orders")); + } + } + + [Test] + public async Task A_known_endpoint_already_in_the_target_keeps_its_flag_and_is_counted_as_already_present() + { + var details = new EndpointDetails { Name = "Sales.Orders", HostId = Guid.NewGuid(), Host = "SALES01" }; + await MonitoringDataStore.CreateIfNotExists(details); + + var endpoint = new KnownEndpoint { EndpointDetails = details, HostDisplayName = "SALES01", Monitored = true }; + var sourceId = $"KnownEndpoints/{details.GetDeterministicId()}"; + + var result = await Target.Write(KnownEndpointsCategory, BatchOf((sourceId, endpoint)), CheckpointAfter(sourceId)); + + using (Assert.EnterMultipleScope()) + { + Assert.That(result.Copied, Is.Zero); + Assert.That(result.AlreadyPresent, Is.EqualTo(1)); + Assert.That((await MonitoringDataStore.GetAllKnownEndpoints()).Single().Monitored, Is.False, "insert-if-absent never updates, so the flag the target already had stays"); + } + } + + [Test] + public async Task A_known_endpoint_with_no_name_or_no_host_is_skipped_as_required_value_missing() + { + var named = new KnownEndpoint { EndpointDetails = new EndpointDetails { Name = "Sales.Orders", HostId = Guid.NewGuid(), Host = "SALES01" } }; + var nameless = new KnownEndpoint { EndpointDetails = new EndpointDetails { Name = null, HostId = Guid.NewGuid(), Host = "SALES02" } }; + var hostless = new KnownEndpoint { EndpointDetails = new EndpointDetails { Name = "Billing", HostId = Guid.NewGuid(), Host = null } }; + + var result = await Target.Write(KnownEndpointsCategory, BatchOf(("KnownEndpoints/1", named), ("KnownEndpoints/2", nameless), ("KnownEndpoints/3", hostless)), CheckpointAfter("KnownEndpoints/3")); + + using (Assert.EnterMultipleScope()) + { + Assert.That(result.Copied, Is.EqualTo(1)); + Assert.That(result.SkippedIds, Is.EquivalentTo(new[] { "KnownEndpoints/2", "KnownEndpoints/3" })); + Assert.That(result.SkipReasons[MigrationSkipReason.RequiredValueMissing], Is.EqualTo(2), "KnownEndpoints.Name and Host are NOT NULL, and a throw here would halt a required category"); + Assert.That((await MonitoringDataStore.GetAllKnownEndpoints()).Select(endpoint => endpoint.EndpointDetails.Name), Is.EqualTo(new[] { "Sales.Orders" })); + } + } + + [Test] + public async Task A_batch_of_nothing_but_skips_saves_the_cursor_and_writes_no_rows() + { + var nameless = new KnownEndpoint { EndpointDetails = new EndpointDetails { Name = null, HostId = Guid.NewGuid(), Host = "SALES02" } }; + var hostless = new KnownEndpoint { EndpointDetails = new EndpointDetails { Name = "Billing", HostId = Guid.NewGuid(), Host = null } }; + + var result = await Target.Write(KnownEndpointsCategory, BatchOf(("KnownEndpoints/1", nameless), ("KnownEndpoints/2", hostless)), CheckpointAfter("KnownEndpoints/2")); + + var stored = await ServiceProvider.GetRequiredService().Read(MigrationCategoryIds.KnownEndpoints); + + using (Assert.EnterMultipleScope()) + { + Assert.That(result.Copied, Is.Zero); + Assert.That(result.AlreadyPresent, Is.Zero, "no key was looked up, so no row may be counted as one the target already held"); + Assert.That(result.Skipped, Is.EqualTo(2)); + Assert.That(await MonitoringDataStore.GetAllKnownEndpoints(), Is.Empty); + Assert.That(stored.Cursor, Is.EqualTo("KnownEndpoints/2"), "a batch that copied nothing still has to commit the cursor past it, or the restart reads the same rows forever"); + } + } + + [Test] + public void A_batch_reporting_more_copied_and_skipped_rows_than_it_held_is_refused() + { + // The guard only counts rows, so what the documents hold cannot change its answer. + var batch = BatchOf(("KnownEndpoints/1", new object()), ("KnownEndpoints/2", new object())); + + var exception = Assert.Throws(() => EFCoreMigrationTarget.AlreadyPresentIn(KnownEndpointsCategory, batch, copied: 2, skipped: 1)); + + Assert.That(exception.Message, Does.Contain("copied 2").And.Contain("skipped 1").And.Contain(MigrationCategoryIds.KnownEndpoints)); + } + + [Test] + public async Task A_batch_whose_checkpoint_save_fails_leaves_no_rows_behind() + { + // The rows save first and the checkpoint second, so this is the only order in which the two can part + // company: a stale version fails the checkpoint after the endpoint row is already in the transaction. + var store = ServiceProvider.GetRequiredService(); + await store.Upsert(CheckpointAfter("KnownEndpoints/0")); + + var endpoint = new KnownEndpoint + { + EndpointDetails = new EndpointDetails { Name = "Sales.Orders", HostId = Guid.NewGuid(), Host = "SALES01" }, + Monitored = true + }; + var sourceId = $"KnownEndpoints/{endpoint.EndpointDetails.GetDeterministicId()}"; + + Assert.ThrowsAsync(async () => + await Target.Write(KnownEndpointsCategory, BatchOf((sourceId, endpoint)), CheckpointAfter(sourceId))); + + using (Assert.EnterMultipleScope()) + { + Assert.That(await MonitoringDataStore.GetAllKnownEndpoints(), Is.Empty, "the rows and the checkpoint commit together, so a failed checkpoint takes the rows with it"); + Assert.That((await store.Read(MigrationCategoryIds.KnownEndpoints)).Cursor, Is.EqualTo("KnownEndpoints/0"), "the stored cursor must still describe the rows the target actually holds"); + } + } + + [Test] + public async Task Every_mapped_column_is_set_from_a_fully_populated_document() + { + var endpoint = new KnownEndpoint + { + EndpointDetails = new EndpointDetails { Name = "Sales.Orders", HostId = Guid.NewGuid(), Host = "SALES01" }, + HostDisplayName = "SALES01", + Monitored = true + }; + var sourceId = $"KnownEndpoints/{endpoint.EndpointDetails.GetDeterministicId()}"; + + await Target.Write(KnownEndpointsCategory, BatchOf((sourceId, endpoint)), CheckpointAfter(sourceId)); + + using var scope = ServiceProvider.CreateScope(); + var dbContext = scope.ServiceProvider.GetRequiredService(); + + MigrationEntityCoverage.AssertEveryMappedPropertyIsSet(dbContext.Model, await dbContext.KnownEndpoints.AsNoTracking().SingleAsync()); + } + + + [Test] + public async Task A_batch_spanning_several_statements_counts_what_each_statement_inserted() + { + var rows = Enumerable.Range(0, 500) + .Select(index => new KnownEndpoint { EndpointDetails = new EndpointDetails { Name = $"Endpoint{index}", HostId = Guid.NewGuid(), Host = "HOST01" }, Monitored = index % 2 == 0 }) + .Select(endpoint => ($"KnownEndpoints/{endpoint.EndpointDetails.GetDeterministicId()}", (object)endpoint)) + .ToArray(); + var batch = BatchOf(rows); + + var first = await Target.Write(KnownEndpointsCategory, batch, CheckpointAfter(batch.Cursor)); + // The engine carries the committed checkpoint into the next write, and the version guard refuses anything else. + var second = await Target.Write(KnownEndpointsCategory, batch, first.Saved with { Cursor = batch.Cursor }); + + using (Assert.EnterMultipleScope()) + { + Assert.That(first.Copied, Is.EqualTo(500)); + Assert.That(second.Copied, Is.Zero, "every key is present on the second write, so no statement may report a row it did not insert"); + Assert.That(second.AlreadyPresent, Is.EqualTo(500)); + } + } + + + [Test] + public void A_batch_whose_writer_neither_prepared_nor_skipped_a_row_is_refused() + { + // The already-present count is a subtraction, so it only catches a writer that over-reports. A row dropped in silence under-reports, and without this guard it is counted as a row the target already held. + var batch = BatchOf(("KnownEndpoints/1", new object()), ("KnownEndpoints/2", new object())); + var prepared = new PreparedBatch((_, _) => Task.FromResult(1), PreparedRowCount: 1, Skips: [], Merges: []); + + var exception = Assert.Throws(() => EFCoreMigrationTarget.AccountForEveryRow(KnownEndpointsCategory, batch, prepared)); + + Assert.That(exception.Message, Does.Contain("prepared 1").And.Contain("skipped 0").And.Contain(MigrationCategoryIds.KnownEndpoints)); + } + + static MigrationBatch BatchOf(params (string Id, object Document)[] rows) => + new([.. rows.Select(row => new MigrationRow(row.Id, row.Document, new Dictionary()))], rows[^1].Id); + + static MigrationCheckpoint CheckpointAfter(string cursor) => + new(MigrationCategoryIds.KnownEndpoints, MigrationCategoryState.InProgress, cursor, 0, 0, null, null, null, null, null, null); +} diff --git a/src/ServiceControl.Persistence.Tests/EFCore/Migration/MigrationCategoryWritersTests.cs b/src/ServiceControl.Persistence.Tests/EFCore/Migration/MigrationCategoryWritersTests.cs new file mode 100644 index 0000000000..7fa6186c1e --- /dev/null +++ b/src/ServiceControl.Persistence.Tests/EFCore/Migration/MigrationCategoryWritersTests.cs @@ -0,0 +1,59 @@ +namespace ServiceControl.Persistence.Tests; + +using System; +using System.Collections.Generic; +using System.Linq; +using System.Threading.Tasks; +using Microsoft.Extensions.DependencyInjection; +using NUnit.Framework; +using ServiceControl.Persistence.DataMigration; +using ServiceControl.Persistence.EFCore.DataMigration; + +class MigrationCategoryWritersTests : PersistenceTestBase +{ + [Test] + public void Every_writer_claims_a_category_the_registry_knows() + { + var unknown = ServiceProvider.GetRequiredService().SupportedCategoryIds + .Where(id => MigrationCategoryRegistry.Find(id) is null) + .ToArray(); + + Assert.That(unknown, Is.Empty, "a writer for an id no category has is dead code the engine can never reach"); + } + + [Test] + public void Every_writer_derives_from_the_typed_base() + { + var writers = typeof(IMigrationCategoryWriter).Assembly.GetTypes() + .Where(type => type is { IsClass: true, IsAbstract: false } && typeof(IMigrationCategoryWriter).IsAssignableFrom(type)) + .ToArray(); + + Assert.That(writers, Is.Not.Empty, "the test proves nothing if it finds no writer"); + Assert.That(writers.Where(type => !DerivesFrom(type, typeof(MigrationCategoryWriter<>))).Select(type => type.Name), Is.Empty, + "a writer that implements the interface directly skips the named cast and can declare a DocumentType it does not take"); + } + + static bool DerivesFrom(Type type, Type genericBase) => + type.BaseType is { } baseType && ((baseType.IsGenericType && baseType.GetGenericTypeDefinition() == genericBase) || DerivesFrom(baseType, genericBase)); + + [Test] + public async Task Every_writer_names_its_category_and_the_row_when_a_document_is_not_its_type() + { + var target = ServiceProvider.GetRequiredService(); + await target.Open(); + + using (Assert.EnterMultipleScope()) + { + foreach (var (categoryId, document) in target.SupportedCategoryIds.SelectMany(id => new (string, object)[] { (id, new object()), (id, null) })) + { + var category = MigrationCategoryRegistry.Find(categoryId)!; + var batch = new MigrationBatch([new MigrationRow("Wrong/1", document, new Dictionary())], "Wrong/1"); + var checkpoint = new MigrationCheckpoint(categoryId, MigrationCategoryState.InProgress, null, 0, 0, null, null, null, null, null, null); + + var exception = Assert.ThrowsAsync(() => target.Write(category, batch, checkpoint)); + + Assert.That(exception!.Message, Does.Contain(categoryId).And.Contain("Wrong/1").And.Contain(target.DocumentTypes[categoryId].FullName), "an engine halt quotes this message, so it has to say which pair disagrees"); + } + } + } +} diff --git a/src/ServiceControl.Persistence.Tests/EFCore/Migration/MigrationEntityCoverage.cs b/src/ServiceControl.Persistence.Tests/EFCore/Migration/MigrationEntityCoverage.cs new file mode 100644 index 0000000000..d3ede8f8ec --- /dev/null +++ b/src/ServiceControl.Persistence.Tests/EFCore/Migration/MigrationEntityCoverage.cs @@ -0,0 +1,34 @@ +namespace ServiceControl.Persistence.Tests; + +using System; +using System.Collections; +using System.Linq; +using Microsoft.EntityFrameworkCore.Metadata; +using NUnit.Framework; + +static class MigrationEntityCoverage +{ + public static void AssertEveryMappedPropertyIsSet(IModel model, TEntity entity, params string[] legitimatelyDefault) + { + var entityType = model.FindEntityType(typeof(TEntity)) + ?? throw new ArgumentException($"{typeof(TEntity).Name} is not an entity in the EF Core model.", nameof(entity)); + + // Reflection rather than a new() constraint, which an entity with required members cannot satisfy. + var defaults = Activator.CreateInstance(typeof(TEntity)); + + // Every property in this model that is not ValueGenerated.Never is an identity column the database assigns. + var unset = entityType.GetProperties() + .Where(property => property.ValueGenerated == ValueGenerated.Never && !legitimatelyDefault.Contains(property.Name)) + .Where(property => SameValue(property.GetGetter().GetClrValue(entity), property.GetGetter().GetClrValue(defaults))) + .Select(property => property.Name) + .ToArray(); + + Assert.That(unset, Is.Empty, $"{typeof(TEntity).Name} has mapped properties equal to a default-constructed {typeof(TEntity).Name}. Set each from the source document, feed the test a document whose values all differ from those defaults, or name the property in legitimatelyDefault if RavenDB genuinely holds that value."); + } + + // Collections compare by item, and an empty one is unset whether or not the entity initialises it. + static bool SameValue(object actual, object defaultValue) => + actual is IEnumerable items and not string + ? !items.Cast().Any() || (defaultValue is IEnumerable defaultItems && items.Cast().SequenceEqual(defaultItems.Cast())) + : Equals(actual, defaultValue); +} diff --git a/src/ServiceControl.Persistence.Tests/EFCore/Migration/MigrationEntityCoverageClaimTests.cs b/src/ServiceControl.Persistence.Tests/EFCore/Migration/MigrationEntityCoverageClaimTests.cs new file mode 100644 index 0000000000..d5d0fedf42 --- /dev/null +++ b/src/ServiceControl.Persistence.Tests/EFCore/Migration/MigrationEntityCoverageClaimTests.cs @@ -0,0 +1,33 @@ +namespace ServiceControl.Persistence.Tests; + +using System.Linq; +using Microsoft.EntityFrameworkCore.Metadata; +using Microsoft.Extensions.DependencyInjection; +using NUnit.Framework; +using ServiceControl.Persistence.EFCore.DbContexts; + +class MigrationEntityCoverageClaimTests : PersistenceTestBase +{ + [Test] + public void Every_property_the_database_fills_in_is_a_key_it_assigns_on_insert() + { + using var scope = ServiceProvider.CreateScope(); + var model = scope.ServiceProvider.GetRequiredService().Model; + + var databaseFilled = model.GetEntityTypes() + .SelectMany(entityType => entityType.GetProperties()) + .Where(property => property.ValueGenerated != ValueGenerated.Never) + .ToArray(); + + var notKeysAssignedOnInsert = databaseFilled + .Where(property => !property.IsPrimaryKey() || property.ValueGenerated != ValueGenerated.OnAdd) + .Select(property => $"{property.DeclaringType.ClrType.Name}.{property.Name} is {property.ValueGenerated}") + .ToArray(); + + using (Assert.EnterMultipleScope()) + { + Assert.That(databaseFilled, Is.Not.Empty, "the model no longer has a single property the database fills in, so this test is watching nothing. Either the identity keys have gone, or the model was never built."); + Assert.That(notKeysAssignedOnInsert, Is.Empty, "MigrationEntityCoverage checks only the properties that are ValueGenerated.Never, on the claim that every other one is a key the database assigns on insert. These properties break that claim, so the coverage check silently stops watching them and a migration can copy a row with them left unset."); + } + } +} diff --git a/src/ServiceControl.Persistence.Tests/EFCore/Migration/MigrationTargetCoverageTests.cs b/src/ServiceControl.Persistence.Tests/EFCore/Migration/MigrationTargetCoverageTests.cs new file mode 100644 index 0000000000..da87ec7dcd --- /dev/null +++ b/src/ServiceControl.Persistence.Tests/EFCore/Migration/MigrationTargetCoverageTests.cs @@ -0,0 +1,26 @@ +namespace ServiceControl.Persistence.Tests; + +using System.Threading; +using System.Threading.Tasks; +using Microsoft.Extensions.DependencyInjection; +using NUnit.Framework; +using ServiceControl.Persistence.DataMigration; + +class MigrationTargetCoverageTests : PersistenceTestBase +{ + [Test] + public async Task Every_category_the_target_supports_has_a_batch_size_and_a_count() + { + var target = ServiceProvider.GetRequiredService(); + + Assert.That(target.SupportedCategoryIds, Is.SupersetOf(new[] { MigrationCategoryIds.KnownEndpoints, MigrationCategoryIds.EndpointSettings }), "the test proves nothing if the target supports no category"); + + foreach (var id in target.SupportedCategoryIds) + { + var category = MigrationCategoryRegistry.Find(id)!; + + Assert.That(await target.BatchSizeFor(category, CancellationToken.None), Is.Positive, $"{id}: every supported category has a batch size"); + Assert.That(await target.Count(category, CancellationToken.None), Is.Zero, id); + } + } +} diff --git a/src/ServiceControl.Persistence.Tests/EFCore/Migration/MigrationTargetReadinessTests.cs b/src/ServiceControl.Persistence.Tests/EFCore/Migration/MigrationTargetReadinessTests.cs new file mode 100644 index 0000000000..caffb35e0e --- /dev/null +++ b/src/ServiceControl.Persistence.Tests/EFCore/Migration/MigrationTargetReadinessTests.cs @@ -0,0 +1,226 @@ +namespace ServiceControl.Persistence.Tests; + +using System; +using System.IO; +using System.Linq; +using System.Threading.Tasks; +using Microsoft.EntityFrameworkCore; +using Microsoft.EntityFrameworkCore.Infrastructure; +using Microsoft.EntityFrameworkCore.Migrations; +using Microsoft.Extensions.DependencyInjection; +using Microsoft.Extensions.Hosting; +using NUnit.Framework; +using ServiceControl.Persistence.DataMigration; +using ServiceControl.Persistence.EFCore.Abstractions; +using ServiceControl.Persistence.EFCore.DataMigration; +using ServiceControl.Persistence.EFCore.DbContexts; +using ServiceControl.Persistence.EFCore.Entities; +using ServiceControl.Persistence.EFCore.Implementation.BodyStorage; +using ServiceControl.Persistence.EFCore.Infrastructure; + +class MigrationTargetReadinessTests : PersistenceTestBase +{ + IMigrationTargetReadiness Readiness => ServiceProvider.GetRequiredService(); + + [Test] + public void The_target_contributes_the_three_checks_only_it_can_make() => + Assert.That( + Readiness.ContributedChecks().Select(check => check.GetType()), + Is.EqualTo(new[] { typeof(SchemaIsCurrentCheck), typeof(TargetHoldsNoServiceControlDataCheck), typeof(BodyStorageIsWritableCheck) })); + + [Test] + public void Every_contributed_check_passes_against_a_migrated_database() + { + using (Assert.EnterMultipleScope()) + { + foreach (var check in Readiness.ContributedChecks()) + { + Assert.DoesNotThrowAsync(() => check.Run(), $"a healthy target must pass every contributed check, and it failed '{check.Name}'"); + } + } + } + + [Test] + public async Task A_target_holding_a_row_is_refused_while_no_checkpoint_exists() + { + await EndpointSettingsStore.UpdateEndpointSettings(new EndpointSettings { Name = "Sales", TrackInstances = true }); + + var exception = Assert.ThrowsAsync(() => TargetCheck.Run()); + + Assert.That(exception.Message, Does.Contain(TableName()).And.Contain("--setup")); + } + + // A plain start writes this row first, so a key someone later ignores fails here before it reaches a customer. + [Test] + public async Task A_settings_row_counts_as_data() + { + await ServiceProvider.GetRequiredService().StoreTrialEndDate(new DateOnly(2030, 1, 1)); + + var exception = Assert.ThrowsAsync(() => TargetCheck.Run()); + + Assert.That(exception.Message, Does.Contain(TableName())); + } + + [Test] + public async Task A_target_with_a_required_checkpoint_row_is_not_judged() + { + await EndpointSettingsStore.UpdateEndpointSettings(new EndpointSettings { Name = "Sales", TrackInstances = true }); + await SeedCheckpoint(MigrationCategoryIds.KnownEndpoints, MigrationCategoryState.Complete); + + Assert.DoesNotThrowAsync(() => TargetCheck.Run()); + } + + [Test] + public async Task A_row_under_an_id_this_build_does_not_know_counts_as_a_started_copy() + { + await EndpointSettingsStore.UpdateEndpointSettings(new EndpointSettings { Name = "Sales", TrackInstances = true }); + await SeedCheckpoint("SomeCategoryFromANewerBuild", MigrationCategoryState.InProgress); + + Assert.DoesNotThrowAsync(() => TargetCheck.Run()); + } + + // --migration-abandon writes this row before any copy has run, so on a database that already served on SQL it must not switch the check off. + [Test] + public async Task An_abandoned_optional_row_alone_does_not_skip_the_check() + { + await EndpointSettingsStore.UpdateEndpointSettings(new EndpointSettings { Name = "Sales", TrackInstances = true }); + await SeedCheckpoint(MigrationCategoryIds.EventLog, MigrationCategoryState.Abandoned); + + var exception = Assert.ThrowsAsync(() => TargetCheck.Run()); + + Assert.That(exception.Message, Does.Contain(TableName())); + } + + [Test] + public void Every_mapped_entity_but_the_checkpoint_is_judged() + { + using var scope = ServiceProvider.CreateScope(); + var dbContext = scope.ServiceProvider.GetRequiredService(); + + var mapped = dbContext.Model.GetEntityTypes().Select(type => type.ClrType).Where(type => type != typeof(MigrationCheckpointEntity)); + + Assert.That(TargetHoldsNoServiceControlDataCheck.Tables.Select(entry => entry.Entity), Is.EquivalentTo(mapped)); + } + + IMigrationStartupCheck TargetCheck => Readiness.ContributedChecks().OfType().Single(); + + Task SeedCheckpoint(string categoryId, MigrationCategoryState state) => + ServiceProvider.GetRequiredService().Upsert(new MigrationCheckpoint(categoryId, state, null, 0, 0, null, null, null, null, null, null)); + + string TableName() + { + using var scope = ServiceProvider.CreateScope(); + + return scope.ServiceProvider.GetRequiredService().Model.FindEntityType(typeof(T))!.GetTableName()!; + } + + [Test] + public async Task The_schema_check_refuses_a_database_whose_migrations_have_not_been_applied() + { + using var scope = ServiceProvider.CreateScope(); + var dbContext = scope.ServiceProvider.GetRequiredService(); + + // EF's own history repository, because it is the only thing that knows where the history table is once the persister is given a schema. + var history = dbContext.GetService(); + + foreach (var row in await history.GetAppliedMigrationsAsync()) + { + await dbContext.Database.ExecuteSqlRawAsync(history.GetDeleteScript(row.MigrationId)); + } + + var exception = Assert.ThrowsAsync(() => Readiness.ContributedChecks().OfType().Single().Run()); + + Assert.That(exception.Message, Does.Contain(dbContext.Database.GetMigrations().First()).And.Contain("--setup")); + } + + [Test] + public async Task The_body_storage_check_leaves_no_probe_body_behind() + { + var bodyStorage = ServiceProvider.GetRequiredService(); + + await Readiness.ContributedChecks().Single(check => check.Name == "message body storage is writable").Run(); + + Assert.That(await bodyStorage.ReadBody("migration-writable-probe"), Is.Null, "a probe left in the store is a body the source never had, which a count comparison reads as an extra rather than a loss"); + } + + [Test] + public void The_body_storage_check_refuses_a_store_that_cannot_write() + { + var parentThatIsAFile = Path.Combine(Path.GetTempPath(), $"sc-not-a-directory-{Guid.NewGuid():n}"); + File.WriteAllText(parentThatIsAFile, string.Empty); + + try + { + var check = new BodyStorageIsWritableCheck(new FileSystemBodyStoragePersistence( + new FileSystemBodyStorageSettings { StoragePath = Path.Combine(parentThatIsAFile, "bodies") })); + + Assert.CatchAsync(() => check.Run()); + } + finally + { + File.Delete(parentThatIsAFile); + } + } + + // The copy runs before any hosted service starts, so a target that only works once one has started is broken exactly when the migration needs it. + [Test] + public async Task The_target_answers_from_a_container_whose_hosted_services_have_not_started() + { + var (host, context) = await BuildHostWithoutStarting(); + + try + { + await using var scope = host.Services.GetRequiredService().CreateAsyncScope(); + + Assert.That(await scope.ServiceProvider.GetRequiredService().EndpointSettings.LongCountAsync(), Is.Zero); + Assert.DoesNotThrowAsync(() => host.Services.GetRequiredService().ContributedChecks().Single(check => check.Name == "message body storage is writable").Run()); + } + finally + { + await context.TearDown(); + host.Dispose(); + } + } + + // PersistenceTestBase already started a host in SetUp, so this builds a second one over its own test database. + static async Task<(IHost Host, PersistenceTestsContext Context)> BuildHostWithoutStarting() + { + var context = new PersistenceTestsContext(); + var hostBuilder = Host.CreateApplicationBuilder(); + + await context.Setup(hostBuilder); + var host = hostBuilder.Build(); + await context.InstallSchema(host); + + return (host, context); + } + + [Test] + public async Task The_host_opened_record_is_absent_until_written_and_stays_after_a_second_write() + { + var readiness = ServiceProvider.GetRequiredService(); + + Assert.That(await readiness.HasHostOpened(), Is.False, "a target nothing has opened on must not claim otherwise"); + + await readiness.RecordHostOpened(); + var firstOpened = Now; + + AdvanceClock(TimeSpan.FromDays(7)); + await readiness.RecordHostOpened(); + + using (Assert.EnterMultipleScope()) + { + Assert.That(await readiness.HasHostOpened(), Is.True); + // The stored instant is what tells an operator the clean abort is over, so a later start must not move it forward. + Assert.That(await ReadHostOpenedAt(), Is.EqualTo(firstOpened)); + } + } + + async Task ReadHostOpenedAt() + { + using var scope = ServiceProvider.CreateScope(); + var dbContext = scope.ServiceProvider.GetRequiredService(); + + return await dbContext.GetSetting(SettingKeys.MigrationHostOpenedOnTarget); + } +} diff --git a/src/ServiceControl.Persistence.Tests/EFCore/MigrationCheckpointTableTests.cs b/src/ServiceControl.Persistence.Tests/EFCore/MigrationCheckpointTableTests.cs index 4d92080e2f..5261d9fffd 100644 --- a/src/ServiceControl.Persistence.Tests/EFCore/MigrationCheckpointTableTests.cs +++ b/src/ServiceControl.Persistence.Tests/EFCore/MigrationCheckpointTableTests.cs @@ -109,6 +109,23 @@ public void The_skip_reasons_JSON_is_keyed_by_reason_name_not_by_its_number() Assert.That(json, Does.Contain("BodyUnreadable")); } + [Test] + public void A_blank_group_comment_skip_round_trips_by_name() + { + using var scope = ServiceProvider.CreateScope(); + var dbContext = scope.ServiceProvider.GetRequiredService(); + var converter = dbContext.Model.FindEntityType(typeof(MigrationCheckpointEntity))!.FindProperty("SkipReasons")!.GetValueConverter()!; + + var json = (string)converter.ConvertToProvider(new Dictionary { [MigrationSkipReason.BlankGroupComment] = 3 })!; + var reasons = (IReadOnlyDictionary)converter.ConvertFromProvider(json)!; + + using (Assert.EnterMultipleScope()) + { + Assert.That(json, Does.Contain("BlankGroupComment")); + Assert.That(reasons, Is.EquivalentTo(new Dictionary { [MigrationSkipReason.BlankGroupComment] = 3 })); + } + } + [Test] public void A_reason_name_this_build_does_not_know_is_read_as_Unknown_rather_than_throwing() { @@ -145,7 +162,7 @@ public void The_table_carries_exactly_the_columns_the_checkpoint_needs() Assert.That(entityType.GetProperties().Select(property => property.Name), Is.EquivalentTo(new[] { "CategoryId", "State", "Cursor", "CopiedCount", "SkippedCount", "SourceTotal", - "SkipReasons", "StartedAt", "LastProgressAt", "SettledAt", "LastError", "AlreadyPresentCount", "Version" + "SkipReasons", "StartedAt", "LastProgressAt", "SettledAt", "LastError", "AlreadyPresentCount", "Version", "StartedWindowSeconds" })); // Underscores stripped so one assertion covers MigrationCheckpoints and migration_checkpoints. diff --git a/src/ServiceControl.Persistence.Tests/IPersistenceTestsContext.cs b/src/ServiceControl.Persistence.Tests/IPersistenceTestsContext.cs index 15d9b76981..5e238f2d22 100644 --- a/src/ServiceControl.Persistence.Tests/IPersistenceTestsContext.cs +++ b/src/ServiceControl.Persistence.Tests/IPersistenceTestsContext.cs @@ -10,6 +10,13 @@ public interface IPersistenceTestsContext { Task Setup(IHostApplicationBuilder hostBuilder); + /// + /// Puts the schema in place on the built host, before it is started, the way --setup does in + /// production. The EF Core persisters refuse to start against a database whose schema predates the + /// build, and that check runs before any hosted service, so migrating after the start is too late. + /// + Task InstallSchema(IHost host); + Task PostSetup(IHost host); Task TearDown(); diff --git a/src/ServiceControl.Persistence.Tests/PersistenceTestBase.cs b/src/ServiceControl.Persistence.Tests/PersistenceTestBase.cs index f065c190f6..dc5ae1178e 100644 --- a/src/ServiceControl.Persistence.Tests/PersistenceTestBase.cs +++ b/src/ServiceControl.Persistence.Tests/PersistenceTestBase.cs @@ -53,6 +53,7 @@ public async Task SetUp() host = hostBuilder.Build(); + await PersistenceTestsContext.InstallSchema(host); await host.StartAsync(); await PersistenceTestsContext.PostSetup(host); } diff --git a/src/ServiceControl.Persistence/DataMigration/CheckpointMigrationState.cs b/src/ServiceControl.Persistence/DataMigration/CheckpointMigrationState.cs new file mode 100644 index 0000000000..7bf02d735b --- /dev/null +++ b/src/ServiceControl.Persistence/DataMigration/CheckpointMigrationState.cs @@ -0,0 +1,42 @@ +namespace ServiceControl.Persistence.DataMigration; + +using System.Collections.Generic; +using System.Linq; +using System.Threading; +using System.Threading.Tasks; + +/// +/// Answers from the saved checkpoints, read once before the host opens rather than queried live. +/// +public sealed class CheckpointMigrationState(IMigrationCheckpointStore? checkpointStore = null) : IMigrationState +{ + volatile bool anyCategoryIncomplete; + IReadOnlyCollection selectedCategoryIds = []; + + public bool AnyCategoryIncomplete => anyCategoryIncomplete; + + /// + /// Reads every checkpoint and fixes the answer for the life of the host. + /// + /// The required categories this instance copies, which are the only ones that hold back the services that delete or overwrite copied data. They are passed in because the store holds a row only for a category some run has already been given. + public async Task Seed(IReadOnlyCollection selectedIds, CancellationToken cancellationToken = default) + { + // Only a persister that can be migrated into keeps checkpoints, so no store means no migration has run. + if (checkpointStore is null) + { + return; + } + + selectedCategoryIds = selectedIds; + + Recompute(await checkpointStore.ReadAll(cancellationToken)); + } + + /// + /// Works the answer out again from the checkpoints given, without reading the store. + /// + /// Every checkpoint the store holds. A selected category with no checkpoint among them has not started, which is not finished. + public void Recompute(IReadOnlyList checkpoints) => + anyCategoryIncomplete = selectedCategoryIds.Any(id => + checkpoints.FirstOrDefault(checkpoint => checkpoint.CategoryId == id)?.State.IsFinished() != true); +} diff --git a/src/ServiceControl.Persistence/DataMigration/HaltThreshold.cs b/src/ServiceControl.Persistence/DataMigration/HaltThreshold.cs index 92d6a73ce7..534e120b7d 100644 --- a/src/ServiceControl.Persistence/DataMigration/HaltThreshold.cs +++ b/src/ServiceControl.Persistence/DataMigration/HaltThreshold.cs @@ -1,8 +1,19 @@ namespace ServiceControl.Persistence.DataMigration; +/// +/// The rule that stops a category early because it has lost too much to keep copying. +/// public static class HaltThreshold { - // Both must be exceeded: the floor ignores a few bad rows in a small category, the percentage a small share of a large one. + /// + /// Whether the skipped rows are enough to stop the category. Both limits have to be passed, so the floor + /// ignores a few bad rows in a small category and the percentage ignores a small share of a large one. + /// + /// Rows skipped as faults. A skip the product would have dropped anyway does not belong here. + /// Rows dealt with over the same stretch as , copied, skipped and already present alike. + /// The share of skipped rows, as a percentage, that has to be passed. + /// The number of skipped rows that has to be passed before the percentage counts at all. + /// True when both limits are passed, which means the category stops. public static bool Exceeded(long skippedCount, long totalCount, int percentThreshold, int minimumFloor) { if (skippedCount <= minimumFloor || totalCount == 0) diff --git a/src/ServiceControl.Persistence/DataMigration/IMigrationCheckpointStore.cs b/src/ServiceControl.Persistence/DataMigration/IMigrationCheckpointStore.cs index bc71283d01..0731525dab 100644 --- a/src/ServiceControl.Persistence/DataMigration/IMigrationCheckpointStore.cs +++ b/src/ServiceControl.Persistence/DataMigration/IMigrationCheckpointStore.cs @@ -6,18 +6,61 @@ namespace ServiceControl.Persistence.DataMigration; using System.Threading; using System.Threading.Tasks; +/// +/// Where one category has got to. says which of these +/// let the host open, and which wait for the operator. A +/// start picks up every category that is neither. +/// public enum MigrationCategoryState { + /// No run has read a row of this category yet. NotStarted, + + /// + /// A run is copying this category, or a run stopped without settling it. On an optional category, a LastError + /// means an exception stopped it or a start could not open the source, and the next start resumes it. + /// InProgress, + + /// + /// The category reached the end of the source with no fault skips. Harmless ones stay counted. + /// Complete, + + /// + /// The category reached the end of the source with fault skips, so it is Failed until the operator retries or abandons it. + /// CompleteWithErrors, + + /// + /// The category stopped early. It waits for --migration-retry or --migration-abandon, and no start re-reads it. + /// Halted, + + /// + /// The operator gave up on a Failed or started required category, on any optional one, or on one whose rows + /// depend on such a category. It is final: nothing copies it again, even in a later migration. + /// Abandoned, + + /// + /// This category follows a category that is not Done or Abandoned, so it did not run. A restart clears it once that one is. + /// Blocked } -/// A category's saved progress: where a restart carries on from, and what the status and verify commands report. +/// +/// A category's saved progress: where a restart carries on from, and what the status and verify commands report. +/// Every count is the total across every run, not this run alone. +/// +/// The point the last committed batch reached. A restart reads the source after it. Null means nothing has been read. +/// No longer written: the engine does not count the source, so this is null on every row it saves. The column stays only because dropping it is a schema change. +/// How many rows each reason skipped. The counts here add up to . +/// The moment the category stopped running, whatever state it stopped in. Read it beside , because a halt settles too. +/// Why the category stopped or did not run, in the words the operator is shown: a halt, a category it must follow that is not finished, an exception that left an optional category copying, or a source a start could not open, on an optional category still copying or not started. It stays until a start runs the category again. +/// Rows the target already held, so they were neither copied nor skipped. They still count as accounted for. +/// The optimistic concurrency token, which is the guard against two writers. A store sets it on save and refuses a checkpoint carrying a value the stored row no longer holds. +/// The window, in whole seconds, an optional category started with. It is null until the copier writes it, and always null for a required category. public sealed record MigrationCheckpoint( string CategoryId, MigrationCategoryState State, @@ -28,14 +71,22 @@ public sealed record MigrationCheckpoint( IReadOnlyDictionary? SkipReasons, DateTime? StartedAt, DateTime? LastProgressAt, - // The moment the category stopped running, whatever state it stopped in. Read it beside State: a halt settles too. DateTime? SettledAt, string? LastError, long AlreadyPresentCount = 0, - // The optimistic concurrency token. A store sets it on save and refuses one carrying a value the stored row no longer holds. - long Version = 0) + long Version = 0, + long? StartedWindowSeconds = null) { - /// Adds one batch's outcome to this checkpoint. A target calls it inside the transaction that writes the rows, so the saved counts are the real ones. + /// + /// Adds one batch's outcome to this checkpoint. A target calls it inside the transaction that writes the rows, + /// so the saved counts are the real ones. + /// + /// Rows this batch wrote. + /// Rows this batch could not write. Every one of them needs a reason. + /// Rows this batch found the target already held. + /// How many rows each reason skipped in this batch. + /// A copy with this batch's counts and reasons added to the totals. + /// The reasons do not add up to , which would leave the verify command unable to account for a row. public MigrationCheckpoint Extend(int copied, int skipped, int alreadyPresent, IReadOnlyDictionary? skipReasons) { var explained = skipReasons?.Values.Sum() ?? 0; @@ -70,11 +121,26 @@ public MigrationCheckpoint Extend(int copied, int skipped, int alreadyPresent, I } } +/// +/// Where the checkpoints are kept. The target database holds them, so progress and the rows it describes +/// commit together. +/// public interface IMigrationCheckpointStore { + /// + /// Every checkpoint the store holds. A run saves a not-started row for every category it will copy before it + /// copies any of them, so a category is absent when no run has recorded it. + /// Task> ReadAll(CancellationToken cancellationToken = default); + + /// + /// One category's checkpoint, or null when no run has recorded it. + /// Task Read(string categoryId, CancellationToken cancellationToken = default); - /// Saves the checkpoint and returns it as stored, carrying the version the save landed on. Throws when the stored row has moved on. + /// + /// Saves the checkpoint and returns it as stored, carrying the version the save landed on. + /// + /// The stored row has moved on, which means another writer saved it, so this save is refused. Task Upsert(MigrationCheckpoint checkpoint, CancellationToken cancellationToken = default); } diff --git a/src/ServiceControl.Persistence/DataMigration/IMigrationSource.cs b/src/ServiceControl.Persistence/DataMigration/IMigrationSource.cs index bd19e623aa..7e969ffa5f 100644 --- a/src/ServiceControl.Persistence/DataMigration/IMigrationSource.cs +++ b/src/ServiceControl.Persistence/DataMigration/IMigrationSource.cs @@ -5,21 +5,41 @@ namespace ServiceControl.Persistence.DataMigration; using System.Threading; using System.Threading.Tasks; +/// +/// Implemented by a persister that can be read as the old database a migration copies from. Nothing here +/// writes to the source, because the old database has to stay usable if the migration is thrown away. +/// public interface IMigrationSource : IAsyncDisposable { - /// Connects to the source read-only. Every other member throws until this has run. + /// + /// Connects to the source read-only. Every other member except throws + /// until this has run. + /// Task Open(CancellationToken cancellationToken = default); - /// What the source report and dry run print about the source. + /// + /// The checks this source wants run before the copy starts, in the order they must run. Call it after Open. + /// + IReadOnlyList ContributedChecks(); + + /// + /// What the source report and the dry run print about the source. + /// Task Describe(CancellationToken cancellationToken = default); - /// A count of everything the source holds, including data no category copies. + /// + /// A count of everything the source holds, including data no category copies. + /// Task> Inventory(CancellationToken cancellationToken = default); - /// How many rows the source holds for one category, for progress and verify. + /// + /// How many rows the source holds for one category, which --migration-verify reports. + /// Task Count(MigrationCategory category, CancellationToken cancellationToken = default); - /// Reads a category in batches, after the checkpoint cursor if provided, or from the start when it is null. Throws on a cursor it never issued. + /// + /// Reads a category in batches, after the checkpoint cursor, or from the start when it is null. Throws on a cursor it never issued. + /// /// A ceiling, not a target: returning fewer costs nothing, returning more fails the target's write. IAsyncEnumerable Read( MigrationCategory category, @@ -27,6 +47,21 @@ IAsyncEnumerable Read( int batchSize, CancellationToken cancellationToken = default); - /// Reads one row's message body when Read did not attach it. Returns null when the row has no body. + /// + /// Reads one row's message body when Read did not attach it. Returns null when the row has no body. + /// Task ReadBody(MigrationCategory category, string sourceId, CancellationToken cancellationToken = default); + + /// + /// The categories this source can read. A category outside this set is never handed to the engine. + /// Answers before Open and does not change across it. + /// + IReadOnlyCollection SupportedCategoryIds { get; } + + /// + /// The type of in every row this source reads, for each category in + /// . The target's writer for the same category must take that type. + /// Answers before Open and does not change across it. + /// + IReadOnlyDictionary DocumentTypes { get; } } diff --git a/src/ServiceControl.Persistence/DataMigration/IMigrationStartupCheck.cs b/src/ServiceControl.Persistence/DataMigration/IMigrationStartupCheck.cs new file mode 100644 index 0000000000..bd08c1df27 --- /dev/null +++ b/src/ServiceControl.Persistence/DataMigration/IMigrationStartupCheck.cs @@ -0,0 +1,21 @@ +namespace ServiceControl.Persistence.DataMigration; + +using System.Threading; +using System.Threading.Tasks; + +/// +/// One check the host runs before the copy starts. The host, the source and the target each contribute their own. +/// +public interface IMigrationStartupCheck +{ + /// + /// What the check is called in the refusal the host prints, so it is phrased to finish the sentence + /// "Migration startup check '...' failed". + /// + string Name { get; } + + /// + /// Runs the check. Throws when it fails, which stops the host before anything is copied. + /// + Task Run(CancellationToken cancellationToken = default); +} diff --git a/src/ServiceControl.Persistence/DataMigration/IMigrationState.cs b/src/ServiceControl.Persistence/DataMigration/IMigrationState.cs new file mode 100644 index 0000000000..212e8694a5 --- /dev/null +++ b/src/ServiceControl.Persistence/DataMigration/IMigrationState.cs @@ -0,0 +1,13 @@ +namespace ServiceControl.Persistence.DataMigration; + +/// +/// Whether this instance is still part way through copying its old database in. +/// +public interface IMigrationState +{ + /// + /// True while any required category is unfinished. Services that delete or overwrite copied data stand down while it + /// is true. Optional categories never make it true, because they copy beside those services. + /// + bool AnyCategoryIncomplete { get; } +} diff --git a/src/ServiceControl.Persistence/DataMigration/IMigrationTarget.cs b/src/ServiceControl.Persistence/DataMigration/IMigrationTarget.cs index 78026fc571..1e457a27a8 100644 --- a/src/ServiceControl.Persistence/DataMigration/IMigrationTarget.cs +++ b/src/ServiceControl.Persistence/DataMigration/IMigrationTarget.cs @@ -1,27 +1,64 @@ namespace ServiceControl.Persistence.DataMigration; +using System; using System.Collections.Generic; using System.Threading; using System.Threading.Tasks; -/// Implemented by a persister that can be the new database a migration copies into. +/// +/// Implemented by a persister that can be the new database a migration copies into. +/// public interface IMigrationTarget { - /// The most rows a source may return in one batch for this category. The target picks it because its own database sets the limit, and a batch over it fails the write. - int BatchSizeFor(MigrationCategory category); + /// + /// Makes the target ready to write. The host calls it before the copy and after the target's own checks have passed. + /// + Task Open(CancellationToken cancellationToken = default); - /// Saves the batch's rows and the checkpoint in one transaction, so progress never gets ahead of the data. Extend checkpointToExtend with this batch's own outcome through and save the result, so what lands is the real split rather than a guess the next save has to correct. - /// Prior totals, the cursor this batch reached, and any rows the engine itself skipped. Not yet counting anything the target does. + /// + /// The most rows a source may return in one batch for this category. + /// The target picks it because its own database sets the limit, and a bigger batch fails the write. + /// + Task BatchSizeFor(MigrationCategory category, CancellationToken cancellationToken = default); + + /// + /// Saves the batch's rows and the checkpoint in one transaction, so progress never gets ahead of the data. + /// Add this batch's own outcome with and save that, so the stored counts are never a guess. + /// + /// Prior totals, the cursor this batch reached, and any rows the engine itself skipped. Nothing the target does is counted yet. Task Write( MigrationCategory category, MigrationBatch batch, MigrationCheckpoint checkpointToExtend, CancellationToken cancellationToken = default); - /// How many rows the target holds for one category, for progress and verify. Counts only that category, even where two categories share a table. + /// + /// How many rows the target holds for one category, for progress and verify. + /// Counts only that category, even where two categories share a table. + /// Task Count(MigrationCategory category, CancellationToken cancellationToken = default); + + /// + /// The categories this target can write. A category outside this set is never handed to the engine. + /// Answers before Open and does not change across it. + /// + IReadOnlyCollection SupportedCategoryIds { get; } + + /// + /// The type of this target takes, for each category in + /// . A row holding any other type fails its batch with an + /// . Answers before Open and does not change across it. + /// + IReadOnlyDictionary DocumentTypes { get; } } -/// What the target did with one batch, and the checkpoint it committed alongside the rows. Every skipped row must have a reason in SkipReasons. -/// How many of Skipped the target would have deleted anyway, such as a row already past retention. Counted and reported like any skip, but never counted toward the halt threshold. -public sealed record MigrationWriteResult(MigrationCheckpoint Saved, int Copied, int Skipped, IReadOnlyList SkippedIds, int AlreadyPresent = 0, IReadOnlyDictionary? SkipReasons = null, int BenignSkipped = 0); +/// +/// What the target did with one batch, and the checkpoint it committed alongside the rows. Every skipped row must have a reason in SkipReasons. +/// +/// The checkpoint as the target stored it, carrying the version that save landed on. The engine carries on from this one, never from the one it passed in. +/// Rows this batch wrote. The same number must show up as the rise in the saved copied count, because the halt threshold reads one and status and verify read the other. +/// Rows this batch could not write, benign ones included. +/// The source ids of those rows, so the engine can name each one in the log. +/// Rows the target already held. They were neither copied nor skipped, and they still count as accounted for. +/// How many rows each reason skipped. The counts have to add up to Skipped, or the checkpoint refuses the batch. +public sealed record MigrationWriteResult(MigrationCheckpoint Saved, int Copied, int Skipped, IReadOnlyList SkippedIds, int AlreadyPresent = 0, IReadOnlyDictionary? SkipReasons = null); diff --git a/src/ServiceControl.Persistence/DataMigration/IMigrationTargetReadiness.cs b/src/ServiceControl.Persistence/DataMigration/IMigrationTargetReadiness.cs new file mode 100644 index 0000000000..8cde500f8f --- /dev/null +++ b/src/ServiceControl.Persistence/DataMigration/IMigrationTargetReadiness.cs @@ -0,0 +1,26 @@ +namespace ServiceControl.Persistence.DataMigration; + +using System.Collections.Generic; +using System.Threading; +using System.Threading.Tasks; + +/// +/// The target's own startup checks, plus a marker kept in the target database saying a host has already started on it. +/// +public interface IMigrationTargetReadiness +{ + /// + /// The checks this target wants run before the copy starts, in the order they must run. + /// + IReadOnlyList ContributedChecks(); + + /// + /// Stamps the marker the first time a host opens on the target, and leaves that first stamp alone on every start after it. + /// + Task RecordHostOpened(CancellationToken cancellationToken = default); + + /// + /// Whether a host has already opened on the target, which is what makes discarding a partial copy a loss rather than a clean abort. + /// + Task HasHostOpened(CancellationToken cancellationToken = default); +} diff --git a/src/ServiceControl.Persistence/DataMigration/MigrationBatch.cs b/src/ServiceControl.Persistence/DataMigration/MigrationBatch.cs index 431730ba82..7a36324b40 100644 --- a/src/ServiceControl.Persistence/DataMigration/MigrationBatch.cs +++ b/src/ServiceControl.Persistence/DataMigration/MigrationBatch.cs @@ -3,9 +3,17 @@ namespace ServiceControl.Persistence.DataMigration; using System; using System.Collections.Generic; +/// +/// When a category is copied. Required categories are copied with ServiceControl closed, because the host must +/// not serve a half copied instance. Optional ones are copied in the background once it is open. +/// public enum MigrationCategoryKind { Required, Optional } /// One kind of data, such as endpoint settings, copied as a unit and resumed from its own cursor. +/// The name in , which is also the key of its checkpoint row. +/// Whether its rows have message bodies, which the engine fetches separately when the source does not attach them. +/// The order within one kind, counting from 1. Required and optional both start at 1, so the two lists are never sorted together. +/// The category that has to finish first, or null when nothing has to. A category whose predecessor is unfinished is recorded as instead of running. public sealed record MigrationCategory( string Id, MigrationCategoryKind Kind, @@ -16,7 +24,11 @@ public sealed record MigrationCategory( /// A message body read from the source, as bytes plus its content type. public sealed record MigrationBody(ReadOnlyMemory Content, string ContentType); -/// One item read from the source. Body is null unless the source attached it or the engine fetched it. +/// One item read from the source. +/// The row's identifier in the old database. It is what the engine logs when the row is skipped, and what takes. +/// The row itself, as the object the source read. The target's writer for that category is what knows the type. +/// Facts about the row that the document does not carry. Empty when the source has none to add. +/// The message body, or null when the source did not attach one and the engine has not fetched it. public sealed record MigrationRow( string SourceId, object Document, @@ -24,4 +36,6 @@ public sealed record MigrationRow( MigrationBody? Body = null); /// Rows read from the source together, plus the cursor to resume after them. +/// The rows, which can be empty when every document in this stretch turned into no row. The cursor past them still has to be saved. +/// Where the source got to. A later read handed this back carries on after it. public sealed record MigrationBatch(IReadOnlyList Rows, string Cursor); diff --git a/src/ServiceControl.Persistence/DataMigration/MigrationCategoryIds.cs b/src/ServiceControl.Persistence/DataMigration/MigrationCategoryIds.cs index 6f15bcdf24..fadc585355 100644 --- a/src/ServiceControl.Persistence/DataMigration/MigrationCategoryIds.cs +++ b/src/ServiceControl.Persistence/DataMigration/MigrationCategoryIds.cs @@ -1,5 +1,9 @@ namespace ServiceControl.Persistence.DataMigration; +/// +/// The name of every category a migration can copy. The name is what the checkpoint row is keyed on, so renaming +/// one strands the progress already saved under the old name. says when each one is copied. +/// public static class MigrationCategoryIds { public const string KnownEndpoints = nameof(KnownEndpoints); @@ -18,7 +22,6 @@ public static class MigrationCategoryIds public const string EventLog = nameof(EventLog); public const string CustomChecks = nameof(CustomChecks); public const string FailedErrorImports = nameof(FailedErrorImports); - public const string FailedMessageEdits = nameof(FailedMessageEdits); public const string ArchivedAndResolvedFailedMessages = nameof(ArchivedAndResolvedFailedMessages); public const string GroupComments = nameof(GroupComments); } diff --git a/src/ServiceControl.Persistence/DataMigration/MigrationCategoryRegistry.cs b/src/ServiceControl.Persistence/DataMigration/MigrationCategoryRegistry.cs index 69665e34e1..13e8a0fbca 100644 --- a/src/ServiceControl.Persistence/DataMigration/MigrationCategoryRegistry.cs +++ b/src/ServiceControl.Persistence/DataMigration/MigrationCategoryRegistry.cs @@ -4,9 +4,23 @@ namespace ServiceControl.Persistence.DataMigration; using System.Linq; using ServiceControl.MessageFailures; +/// +/// Every category a migration can copy, and when each one is copied. This list is the whole of what a migration +/// covers: a source or a target that cannot handle a category says so through its own supported ids, and the +/// engine never invents one. +/// public static class MigrationCategoryRegistry { + /// + /// The failed message statuses that are copied before the host opens. A message in one of these is waiting on + /// somebody, so an operator who cannot see it after the cutover has lost work. + /// public static readonly IReadOnlyList UnresolvedAndRetryIssuedStatuses = [FailedMessageStatus.Unresolved, FailedMessageStatus.RetryIssued]; + + /// + /// The failed message statuses that are copied in the background. These are history, so the instance is usable + /// while they are still arriving. + /// public static readonly IReadOnlyList ArchivedAndResolvedStatuses = [FailedMessageStatus.Archived, FailedMessageStatus.Resolved]; public static readonly IReadOnlyList All = @@ -24,14 +38,13 @@ public static class MigrationCategoryRegistry new(MigrationCategoryIds.LicensingReportMasks, MigrationCategoryKind.Required, CarriesBodies: false, Order: 10), new(MigrationCategoryIds.LicensedEndpointDetails, MigrationCategoryKind.Required, CarriesBodies: false, Order: 11), new(MigrationCategoryIds.UnresolvedAndRetryIssuedFailedMessages, MigrationCategoryKind.Required, CarriesBodies: true, Order: 12), + new(MigrationCategoryIds.CustomChecks, MigrationCategoryKind.Required, CarriesBodies: false, Order: 14), + new(MigrationCategoryIds.FailedErrorImports, MigrationCategoryKind.Required, CarriesBodies: true, Order: 15), + new(MigrationCategoryIds.GroupComments, MigrationCategoryKind.Required, CarriesBodies: false, Order: 16, MustFollow: MigrationCategoryIds.UnresolvedAndRetryIssuedFailedMessages), // Optional, copied in the background once the host is open. new(MigrationCategoryIds.EventLog, MigrationCategoryKind.Optional, CarriesBodies: false, Order: 1), - new(MigrationCategoryIds.CustomChecks, MigrationCategoryKind.Optional, CarriesBodies: false, Order: 2), - new(MigrationCategoryIds.FailedErrorImports, MigrationCategoryKind.Optional, CarriesBodies: true, Order: 3), - new(MigrationCategoryIds.FailedMessageEdits, MigrationCategoryKind.Optional, CarriesBodies: false, Order: 4), - new(MigrationCategoryIds.ArchivedAndResolvedFailedMessages, MigrationCategoryKind.Optional, CarriesBodies: true, Order: 5), - new(MigrationCategoryIds.GroupComments, MigrationCategoryKind.Optional, CarriesBodies: false, Order: 6, MustFollow: MigrationCategoryIds.ArchivedAndResolvedFailedMessages), + new(MigrationCategoryIds.ArchivedAndResolvedFailedMessages, MigrationCategoryKind.Optional, CarriesBodies: true, Order: 2), ]; public static MigrationCategory? Find(string id) => All.FirstOrDefault(c => c.Id == id); diff --git a/src/ServiceControl.Persistence/DataMigration/MigrationCategoryStateExtensions.cs b/src/ServiceControl.Persistence/DataMigration/MigrationCategoryStateExtensions.cs new file mode 100644 index 0000000000..0a6ae5bb2c --- /dev/null +++ b/src/ServiceControl.Persistence/DataMigration/MigrationCategoryStateExtensions.cs @@ -0,0 +1,39 @@ +namespace ServiceControl.Persistence.DataMigration; + +public static class MigrationCategoryStateExtensions +{ + /// + /// Whether the category is done with, so a restart passes over it and it no longer holds the host back: Done, + /// or Abandoned by the operator. A Failed category () is not finished. It waits for + /// --migration-retry or --migration-abandon, and a category that must follow it stays blocked until then. + /// + public static bool IsFinished(this MigrationCategoryState state) => + state is MigrationCategoryState.Complete + or MigrationCategoryState.Abandoned; + + /// + /// Whether the category is Failed: it stopped early, or reached the end having skipped rows to faults. It waits + /// for the operator to run --migration-retry or --migration-abandon, and no start copies it until then. + /// + public static bool IsFailed(this MigrationCategoryState state) => + state is MigrationCategoryState.Halted + or MigrationCategoryState.CompleteWithErrors; +} + +public static class MigrationSkipReasonExtensions +{ + /// + /// Whether the running product would have dropped this row anyway, which makes the skip no loss. The halt + /// threshold and the settle rule never count a harmless skip, so no number of them can stop a category or + /// leave it Failed. The list is fixed. is not on it, because a + /// newer build's reason read back here may be a real loss. + /// + public static bool IsBenign(this MigrationSkipReason reason) => + reason is MigrationSkipReason.PastRetention + or MigrationSkipReason.BlankGroupComment; + + /// + /// Whether no retry can fix a row skipped for this reason, because the row itself can never be stored. + /// + public static bool IsPermanent(this MigrationSkipReason reason) => reason is MigrationSkipReason.RequiredValueMissing; +} diff --git a/src/ServiceControl.Persistence/DataMigration/MigrationEngine.cs b/src/ServiceControl.Persistence/DataMigration/MigrationEngine.cs index 3503321796..bf975e1e21 100644 --- a/src/ServiceControl.Persistence/DataMigration/MigrationEngine.cs +++ b/src/ServiceControl.Persistence/DataMigration/MigrationEngine.cs @@ -7,6 +7,15 @@ namespace ServiceControl.Persistence.DataMigration; using System.Threading.Tasks; using Microsoft.Extensions.Logging; +/// +/// Copies one category at a time from the source to the target, and records where it got to after every batch. +/// The engine knows nothing about either database: what a row is, how it is written and how big a batch can be +/// all come from the source and the target. A category that fails comes back as a Failed checkpoint rather than +/// an exception, and no later run copies it until the operator retries or abandons it. An optional category an +/// exception stopped comes back still copying instead, with the error on its row. A shutdown and a checkpoint +/// conflict come out as exceptions, because neither is the category's fault, and so does a failure to save the +/// checkpoint a category stopped or settled on, because then there is no row to record it on. +/// public sealed class MigrationEngine( IMigrationSource source, IMigrationTarget target, @@ -15,8 +24,13 @@ public sealed class MigrationEngine( MigrationEngineOptions options, ILogger logger) { + /// How many times a message body is read before the row is skipped as unreadable. public const int MaxBodyReadAttempts = 3; + /// + /// The categories of one kind to copy, in the order to copy them. Optional categories the operator did not + /// ask for are left out. + /// public IReadOnlyList SelectCategories(MigrationCategoryKind kind) => MigrationCategoryRegistry.All .Where(c => c.Kind == kind) @@ -24,12 +38,29 @@ public IReadOnlyList SelectCategories(MigrationCategoryKind k .OrderBy(c => c.Order) .ToArray(); + /// + /// Copies the categories one after another and returns where each one ended, in the same order. A category + /// that halts does not stop the ones after it. Before copying anything it saves a not-started checkpoint for + /// every category that has none, so a copy stopped between two categories still lists the ones it never reached. + /// + /// Another writer saved one of these checkpoints, which means a second instance is copying into the same database. + /// The host is shutting down. + /// Saving the checkpoint a category stopped or settled on failed. Its stored row is the last one saved, and the categories after it were not run. // Runs in the order given without re-sorting: required and optional orders both start at 1, so // sorting a mixed list would put an optional category in front of a required one. public async Task> RunCategories( IReadOnlyList categories, CancellationToken cancellationToken = default) { + // The gates that keep a host off an unfinished copy read only the rows that exist. + foreach (var category in categories) + { + if (await checkpointStore.Read(category.Id, cancellationToken) is null) + { + await checkpointStore.Upsert(NotStarted(category), cancellationToken); + } + } + var results = new List(categories.Count); foreach (var category in categories) @@ -40,12 +71,19 @@ public async Task> RunCategories( return results; } + /// + /// Copies one category, carrying on from its saved cursor, and returns the checkpoint it ended on. A category + /// already finished or Failed is returned untouched without reading the source. + /// + /// The stored checkpoint, whose state says how it ended and whose LastError says why it stopped. + /// Another writer saved this category's checkpoint, which means a second instance is copying into the same database. The state is left as that writer set it. + /// The host is shutting down. The last committed batch saved its own counts, so the stored checkpoint is already correct and a restart carries on from it. + /// Saving the checkpoint the category stopped or settled on failed, such as a halt the store could not write. The stored row is the last one saved, so it still reads as copying from its last committed batch. public async Task RunCategoryAsync(MigrationCategory category, CancellationToken cancellationToken = default) { var checkpoint = await checkpointStore.Read(category.Id, cancellationToken) ?? NotStarted(category); - // Halted is deliberately not one of them: a halt says "stopped, and here is why", and a restart after the cause is fixed has to be able to pick it up again. - if (checkpoint.State is MigrationCategoryState.Complete or MigrationCategoryState.CompleteWithErrors or MigrationCategoryState.Abandoned) + if (checkpoint.State.IsFinished() || checkpoint.State.IsFailed()) { return checkpoint; } @@ -53,7 +91,7 @@ public async Task RunCategoryAsync(MigrationCategory catego if (category.MustFollow is { } mustFollowId) { var predecessor = await checkpointStore.Read(mustFollowId, cancellationToken); - if (predecessor is not { State: MigrationCategoryState.Complete or MigrationCategoryState.CompleteWithErrors or MigrationCategoryState.Abandoned }) + if (predecessor?.State.IsFinished() != true) { var predecessorState = predecessor?.State.ToString() ?? "not started"; var blocked = checkpoint with @@ -61,32 +99,40 @@ public async Task RunCategoryAsync(MigrationCategory catego State = MigrationCategoryState.Blocked, LastError = $"Blocked: {category.Id} must follow {mustFollowId}, which is {predecessorState}" }; + logger.LogWarning("Category {CategoryId} did not run: it must follow {PredecessorId}, which is {PredecessorState}", category.Id, mustFollowId, predecessorState); + return await checkpointStore.Upsert(blocked, cancellationToken); } } - if (checkpoint.State is MigrationCategoryState.NotStarted or MigrationCategoryState.Halted or MigrationCategoryState.Blocked) + // A LastError on a row still copying is cleared here, because this start is trying it again. + if (checkpoint.State is MigrationCategoryState.NotStarted or MigrationCategoryState.Blocked || checkpoint.LastError is not null) { checkpoint = await checkpointStore.Upsert(checkpoint with { State = MigrationCategoryState.InProgress, StartedAt = checkpoint.StartedAt ?? timeProvider.GetUtcNow().UtcDateTime, + LastProgressAt = timeProvider.GetUtcNow().UtcDateTime, SettledAt = null, LastError = null }, cancellationToken); } var isFirstBatch = true; - // Per run, not the persisted totals: the skips that tripped a halt stay on the row, so - // counting them again would re-halt a restart whose cause has been fixed. + // Per run, not the persisted totals: the skips an earlier run made stay on the row, so counting them + // again would stop a run resumed from its cursor on its first batch. var skippedThisRun = 0L; var processedThisRun = 0L; + var rowsReadThisRun = 0L; + // Saved after the try rather than inside it, so a failed save is not caught below as the category's own error. + MigrationCheckpoint? halt = null; try { - var batchSize = target.BatchSizeFor(category); + var batchSize = await target.BatchSizeFor(category, cancellationToken); + await foreach (var batch in source.Read(category, checkpoint.Cursor, batchSize, cancellationToken).WithCancellation(cancellationToken)) { if (!isFirstBatch && category.Kind == MigrationCategoryKind.Optional) @@ -94,10 +140,11 @@ public async Task RunCategoryAsync(MigrationCategory catego await Pause(options.ThrottlePause, cancellationToken); } isFirstBatch = false; + rowsReadThisRun += batch.Rows.Count; var batchToWrite = batch; - // Stays off checkpoint until the write commits: the catch persists checkpoint, and a restart - // re-reads an uncommitted batch and would count these skips again. + // Kept out of checkpoint until the write commits: the catch below saves checkpoint, and a + // restart re-reads the uncommitted batch and would count these skips twice. var bodySkips = 0; if (category.CarriesBodies) { @@ -111,43 +158,52 @@ public async Task RunCategoryAsync(MigrationCategory catego } } - // Prior totals, the new cursor, and the rows this engine already skipped. The target adds its own - // outcome inside the transaction that writes the rows, so nothing provisional is ever stored. + // The target adds its own outcome inside the transaction that writes the rows, + // so nothing provisional is ever stored. var checkpointToExtend = checkpoint with { Cursor = batch.Cursor, + LastProgressAt = timeProvider.GetUtcNow().UtcDateTime, SkippedCount = checkpoint.SkippedCount + bodySkips, SkipReasons = MigrationCheckpoint.AddSkipReasons(checkpoint.SkipReasons, bodySkips == 0 ? null : new Dictionary { [MigrationSkipReason.BodyUnreadable] = bodySkips }) }; var result = await target.Write(category, batchToWrite, checkpointToExtend, cancellationToken); + + // The target has committed this row, so every halt below settles from it. A halt settling from the + // older version would be refused by the store as a conflict and lose its reason. checkpoint = result.Saved; - foreach (var id in result.SkippedIds) + // The result states the batch's outcome twice, as its own counts and as deltas on the checkpoint + // it committed. The halt threshold reads the first and status and verify read the second. + var committed = (result.Saved.CopiedCount - checkpointToExtend.CopiedCount, result.Saved.SkippedCount - checkpointToExtend.SkippedCount, result.Saved.AlreadyPresentCount - checkpointToExtend.AlreadyPresentCount); + if (committed != (result.Copied, result.Skipped, result.AlreadyPresent)) { - logger.LogWarning("Skipped {SourceId} in category {CategoryId}", id, category.Id); + var mismatch = $"The target reported copying {result.Copied}, skipping {result.Skipped} and finding {result.AlreadyPresent} already present in category {category.Id}, but the checkpoint it committed moved by {committed}."; + logger.LogError("Category {CategoryId} halted at cursor {Cursor}: {LastError}", category.Id, checkpoint.Cursor, mismatch); + halt = checkpoint with { State = MigrationCategoryState.Halted, LastError = mismatch }; + break; } - // A negative fault count would silently disarm the halt threshold for the rest of the run. - if (result.BenignSkipped > result.Skipped) + foreach (var id in result.SkippedIds) { - throw new InvalidOperationException($"The target reported {result.BenignSkipped} benign skips in category {category.Id} out of {result.Skipped} skipped rows. Benign skips are a subset of the skipped rows."); + logger.LogWarning("Skipped {SourceId} in category {CategoryId}", id, category.Id); } processedThisRun += bodySkips + result.Copied + result.Skipped + result.AlreadyPresent; - // Rows the target would have deleted anyway are not faults, so they never halt a category. - skippedThisRun += bodySkips + result.Skipped - result.BenignSkipped; + skippedThisRun += bodySkips + FaultSkips(result.SkipReasons); if (HaltThreshold.Exceeded(skippedThisRun, processedThisRun, options.HaltThresholdPercent, options.HaltThresholdMinimum)) { - var reason = $"Halted: {skippedThisRun} of {processedThisRun} rows skipped in this run exceeds the configured threshold of {options.HaltThresholdPercent}% and {options.HaltThresholdMinimum} rows. Fix the cause and restart to resume from the cursor, or abandon the category to accept the loss."; + var reason = $"Halted: {skippedThisRun} of {processedThisRun} rows skipped in this run exceeds the configured threshold of {options.HaltThresholdPercent}% and {options.HaltThresholdMinimum} rows."; logger.LogError("Category {CategoryId} halted at cursor {Cursor}: {LastError}", category.Id, checkpoint.Cursor, reason); - return await Settle(checkpoint with { State = MigrationCategoryState.Halted, LastError = reason }, cancellationToken); + halt = checkpoint with { State = MigrationCategoryState.Halted, LastError = reason }; + break; } } } - // A shutdown is not a halt, and there is nothing to reconcile: the last committed batch stored its - // real split with its own rows, so the row on disk is already correct and resumable. + // A shutdown is not a halt: the last committed batch saved its counts with its own rows, + // so the checkpoint on disk is already correct and resumable. catch (OperationCanceledException) when (cancellationToken.IsCancellationRequested) { throw; @@ -161,13 +217,38 @@ public async Task RunCategoryAsync(MigrationCategory catego { var position = checkpoint.Cursor is null ? "at the start" : $"at cursor {checkpoint.Cursor}"; var reason = $"{ex.GetType().Name} {position}: {ex.Message}"; + + // An optional category copies while ServiceControl serves, so an error there waits for the next start rather than for the operator. + if (category.Kind == MigrationCategoryKind.Optional) + { + logger.LogError(ex, "Optional category {CategoryId} stopped at cursor {Cursor} and stays copying until the next start", category.Id, checkpoint.Cursor); + return await checkpointStore.Upsert(checkpoint with { LastError = $"{reason.TrimEnd('.')}. It stays copying and the next start resumes it from the cursor." }, cancellationToken); + } + logger.LogError(ex, "Category {CategoryId} halted at cursor {Cursor}", category.Id, checkpoint.Cursor); return await Settle(checkpoint with { State = MigrationCategoryState.Halted, LastError = reason }, cancellationToken); } - return await Settle(checkpoint with { State = checkpoint.SkippedCount > 0 ? MigrationCategoryState.CompleteWithErrors : MigrationCategoryState.Complete }, cancellationToken); + if (halt is not null) + { + return await Settle(halt, cancellationToken); + } + + // The one check a category runs on itself, judged over this run because the counts on a resumed row cover earlier runs. + if (rowsReadThisRun != processedThisRun) + { + var reason = $"Halted: {category.Id} read {rowsReadThisRun} rows in this run but the target accounted for {processedThisRun} of them as copied, skipped or already present."; + logger.LogError("Category {CategoryId} halted at cursor {Cursor}: {LastError}", category.Id, checkpoint.Cursor, reason); + return await Settle(checkpoint with { State = MigrationCategoryState.Halted, LastError = reason }, cancellationToken); + } + + return await Settle(checkpoint with { State = FaultSkips(checkpoint.SkipReasons) > 0 ? MigrationCategoryState.CompleteWithErrors : MigrationCategoryState.Complete }, cancellationToken); } + // Harmless skips are rows the product would have removed anyway, so they neither stop a category nor leave it Failed. + static long FaultSkips(IReadOnlyDictionary? skipReasons) => + skipReasons is null ? 0 : skipReasons.Where(reason => !reason.Key.IsBenign()).Sum(reason => reason.Value); + // Halts log before settling: the store shares the target's database, so a failed save would hide the cause. Task Settle(MigrationCheckpoint settled, CancellationToken cancellationToken) => checkpointStore.Upsert(settled with { SettledAt = timeProvider.GetUtcNow().UtcDateTime }, cancellationToken); @@ -230,8 +311,8 @@ Task Settle(MigrationCheckpoint settled, CancellationToken static bool IsDefect(Exception exception) => exception is NotSupportedException or NotImplementedException or InvalidOperationException or ArgumentException or NullReferenceException or InvalidCastException; - // A configured pause of zero means "do not throttle", and a timer that is never going to be - // waited on is worse than no timer: against a fake clock nobody advances, it never completes. + // Zero means no throttling, and it must return without waiting: against a fake clock + // nobody advances, even a zero-length wait never finishes. Task Pause(TimeSpan duration, CancellationToken cancellationToken) => duration <= TimeSpan.Zero ? Task.CompletedTask : Task.Delay(duration, timeProvider, cancellationToken); diff --git a/src/ServiceControl.Persistence/DataMigration/MigrationEngineOptions.cs b/src/ServiceControl.Persistence/DataMigration/MigrationEngineOptions.cs index 8433836893..8dee699d9f 100644 --- a/src/ServiceControl.Persistence/DataMigration/MigrationEngineOptions.cs +++ b/src/ServiceControl.Persistence/DataMigration/MigrationEngineOptions.cs @@ -2,10 +2,16 @@ namespace ServiceControl.Persistence.DataMigration; using System; using System.Collections.Generic; -using System.Linq; using ServiceControl.Configuration; -/// The engine's tuning: the pause between optional batches, when a category halts, and which optional categories to copy. +/// +/// The engine's tuning: the pause between optional batches, when a category halts, and which optional categories +/// to copy and how far back. +/// +/// How long to wait between batches of an optional category, which is what keeps the background copy off the instance's back. Zero waits not at all. +/// The share of skipped rows, as a percentage, that stops a category. +/// The number of skipped rows that has to be passed before the percentage counts, so a handful of bad rows in a small category is not a halt. +/// The optional categories to copy. An optional category outside this list is never copied. public sealed record MigrationEngineOptions( TimeSpan ThrottlePause, int HaltThresholdPercent, @@ -17,28 +23,68 @@ public sealed record MigrationEngineOptions( /// How long to wait before trying again to read a message body that failed. Not read from settings. public TimeSpan BodyRetryBackoff { get; init; } = DefaultBodyRetryBackoff; - /// Reads the options from settings. Throws if the optional categories setting names one that doesn't exist or isn't optional. - public static MigrationEngineOptions FromSettings(SettingsRootNamespace settingsRootNamespace) + /// + /// How far back the event log copy goes, measured against when each event was raised. + /// + public TimeSpan EventLogWindow { get; init; } + + /// + /// How far back the archived and resolved failed messages copy goes, measured against when each message was + /// archived or resolved. + /// + public TimeSpan ArchivedAndResolvedFailedMessagesWindow { get; init; } + + /// + /// Reads the options from the instance's settings. Anything not set takes the default from + /// , except the two windows, which default to the retention periods passed in. + /// An optional category is selected when its window is greater than zero. + /// + /// The settings root the migration keys are read under. + /// The instance's event retention period, used as the event log window when none is set. + /// The instance's error retention period, used as the archived and resolved window when none is set. + /// A window setting is not a time span of zero or more, which the operator has to see before the copy starts. + public static MigrationEngineOptions FromSettings(SettingsRootNamespace settingsRootNamespace, TimeSpan eventRetentionPeriod, TimeSpan errorRetentionPeriod) { var throttleMilliseconds = SettingsReader.Read(settingsRootNamespace, MigrationSettings.ThrottlePauseMillisecondsKey, MigrationSettings.DefaultThrottlePauseMilliseconds); var haltPercent = SettingsReader.Read(settingsRootNamespace, MigrationSettings.HaltThresholdPercentKey, MigrationSettings.DefaultHaltThresholdPercent); var haltMinimum = SettingsReader.Read(settingsRootNamespace, MigrationSettings.HaltThresholdMinimumKey, MigrationSettings.DefaultHaltThresholdMinimum); - var optionalCategories = SettingsReader.Read(settingsRootNamespace, MigrationSettings.OptionalCategoriesKey, string.Empty); + var eventLogWindow = ReadWindow(settingsRootNamespace, MigrationSettings.EventLogWindowKey, eventRetentionPeriod); + var archivedAndResolvedWindow = ReadWindow(settingsRootNamespace, MigrationSettings.ArchivedAndResolvedFailedMessagesWindowKey, errorRetentionPeriod); + + var selectedIds = new List(); + + if (eventLogWindow > TimeSpan.Zero) + { + selectedIds.Add(MigrationCategoryIds.EventLog); + } - var selectedIds = optionalCategories - .Split(',', StringSplitOptions.RemoveEmptyEntries | StringSplitOptions.TrimEntries) - .ToArray(); + if (archivedAndResolvedWindow > TimeSpan.Zero) + { + selectedIds.Add(MigrationCategoryIds.ArchivedAndResolvedFailedMessages); + } - var unknown = selectedIds - .Where(id => MigrationCategoryRegistry.Find(id) is not { Kind: MigrationCategoryKind.Optional }) - .ToArray(); + return new MigrationEngineOptions(TimeSpan.FromMilliseconds(throttleMilliseconds), haltPercent, haltMinimum, selectedIds) + { + EventLogWindow = eventLogWindow, + ArchivedAndResolvedFailedMessagesWindow = archivedAndResolvedWindow + }; + } + + static TimeSpan ReadWindow(SettingsRootNamespace settingsRootNamespace, string key, TimeSpan defaultWindow) + { + var value = SettingsReader.Read(settingsRootNamespace, key); + + if (value is null) + { + return defaultWindow; + } - if (unknown.Length > 0) + if (!TimeSpan.TryParse(value, out var window) || window < TimeSpan.Zero) { throw new InvalidOperationException( - $"{MigrationSettings.OptionalCategoriesKey} names categories that do not exist or are not optional: {string.Join(", ", unknown)}"); + $"{key} is '{value}', which is not a time span of zero or more. Set it like a retention period, such as 7.00:00:00 for seven days, or to 0 to leave that category behind."); } - return new MigrationEngineOptions(TimeSpan.FromMilliseconds(throttleMilliseconds), haltPercent, haltMinimum, selectedIds); + return window; } } diff --git a/src/ServiceControl.Persistence/DataMigration/MigrationSettings.cs b/src/ServiceControl.Persistence/DataMigration/MigrationSettings.cs index 0de8ddb892..6102b22058 100644 --- a/src/ServiceControl.Persistence/DataMigration/MigrationSettings.cs +++ b/src/ServiceControl.Persistence/DataMigration/MigrationSettings.cs @@ -1,15 +1,31 @@ namespace ServiceControl.Persistence.DataMigration; -/// The migration's setting names, relative to the instance's settings root, and their defaults. explains the tuning ones. +/// +/// The migration's setting names, relative to the instance's settings root, and their defaults. +/// explains the tuning ones. +/// public static class MigrationSettings { public const string ThrottlePauseMillisecondsKey = "Migration/ThrottlePauseMilliseconds"; public const string HaltThresholdPercentKey = "Migration/HaltThresholdPercent"; public const string HaltThresholdMinimumKey = "Migration/HaltThresholdMinimum"; - /// A comma-separated list of the optional categories to copy, such as "EventLog, CustomChecks". - public const string OptionalCategoriesKey = "Migration/OptionalCategories"; + /// + /// How far back the event log copy goes, as a time span. Unset means the event retention period, and zero + /// turns the category off. + /// + public const string EventLogWindowKey = "Migration/EventLogWindow"; + /// + /// How far back the archived and resolved failed messages copy goes, as a time span. Unset means the error + /// retention period, and zero turns the category off. + /// + public const string ArchivedAndResolvedFailedMessagesWindowKey = "Migration/ArchivedAndResolvedFailedMessagesWindow"; + /// + /// Turns the migration on: the required copy runs before the host opens. + /// + public const string EnabledKey = "Migration/Enabled"; public const int DefaultThrottlePauseMilliseconds = 100; public const int DefaultHaltThresholdPercent = 5; public const int DefaultHaltThresholdMinimum = 100; + public const bool DefaultEnabled = false; } diff --git a/src/ServiceControl.Persistence/DataMigration/MigrationSkipReason.cs b/src/ServiceControl.Persistence/DataMigration/MigrationSkipReason.cs index 4129eb939d..11d03ffa52 100644 --- a/src/ServiceControl.Persistence/DataMigration/MigrationSkipReason.cs +++ b/src/ServiceControl.Persistence/DataMigration/MigrationSkipReason.cs @@ -1,9 +1,30 @@ namespace ServiceControl.Persistence.DataMigration; +/// +/// Why one row was not copied. Every skipped row carries one, so the verify command can account for the +/// difference between the two databases. says which of +/// these are no loss. +/// public enum MigrationSkipReason { + /// The message body could not be read from the source, after the engine had tried more than once. BodyUnreadable, + + /// The row is older than the retention period, so the instance would have deleted it soon anyway. PastRetention, - // Never written by a copier. A database a newer build wrote still reads rather than throwing where the host decides whether to start. + + /// The source row has no value for something the target column requires. + RequiredValueMissing, + + /// + /// The group comment is null or blank, and the product keeps no row for a blank comment. + /// + BlankGroupComment, + + /// + /// A reason a newer build wrote and this one cannot name, so an older host still starts instead of throwing. + /// Nothing in this build writes it. It counts as a fault, because harmless would let this build settle a + /// category Done over rows the newer build lost; a false Failed from it is cleared by one retry. + /// Unknown } diff --git a/src/ServiceControl.Persistence/PersistenceSettings.cs b/src/ServiceControl.Persistence/PersistenceSettings.cs index 08a73eb1c3..4d87e98f76 100644 --- a/src/ServiceControl.Persistence/PersistenceSettings.cs +++ b/src/ServiceControl.Persistence/PersistenceSettings.cs @@ -8,25 +8,44 @@ namespace ServiceControl.Persistence /// public abstract class PersistenceSettings { + /// + /// Whether this host starts the database and nothing else, so an operator can repair or inspect it. + /// Only the RavenDB persister supports it. There the embedded server and RavenDB Studio start as + /// usual, and the host registers nothing that ingests or serves data. On any other persister the + /// host refuses to start in maintenance mode. + /// public bool MaintenanceMode { get; set; } - //HINT: This needs to be here so that ServerControl instance can add an instance specific metadata to tweak the DatabasePath value + + /// + /// The directory that holds the database files, or null when this host holds none. The RavenDB + /// persister fills it in from its DbPath setting, and its disk space checks measure the drive that + /// this path names. The SQL persisters leave it null, because their files live on the database server. + /// public string? DatabasePath { get; set; } /// - /// Whether this host owns the background deletion of data past its retention period. Only one - /// host in a deployment should, so error ingestion only hosts turn it off. + /// Whether this host deletes data that is past its retention period. Only one host in a deployment + /// must do this, so a host that only ingests errors turns it off. The SQL persisters start a + /// background sweeper when it is true. RavenDB expires documents on its own and does not read this. /// public bool RunRetentionSweep { get; set; } = true; + /// + /// Whether message bodies are indexed for search. The RavenDB persister decides this as it ingests a + /// message, so a change only reaches messages ingested after it. The SQL persisters always index the + /// body, so the value does not change what they do. + /// public bool EnableFullTextSearchOnBodies { get; set; } = true; /// - /// Wall-clock limit for the message view queries, see . + /// The wall clock limit for a message query, see . It is also how long + /// this instance waits for a remote instance to answer. /// public TimeSpan QueryTimeout { get; set; } = QueryTimeLimit.Default; /// - /// The setting is read from, as named in the timeout error. + /// The name of the setting that comes from. The timeout error names it, so + /// an operator can see which setting to change. /// public const string QueryTimeoutSettingName = "ServiceControl/" + QueryTimeLimit.SettingName; } diff --git a/src/ServiceControl.UnitTests/ApprovalFiles/APIApprovals.PlatformSampleSettings.approved.txt b/src/ServiceControl.UnitTests/ApprovalFiles/APIApprovals.PlatformSampleSettings.approved.txt index abd4c98313..8a83f80ff0 100644 --- a/src/ServiceControl.UnitTests/ApprovalFiles/APIApprovals.PlatformSampleSettings.approved.txt +++ b/src/ServiceControl.UnitTests/ApprovalFiles/APIApprovals.PlatformSampleSettings.approved.txt @@ -63,6 +63,7 @@ "VirtualDirectory": "", "HeartbeatGracePeriod": "00:00:40", "TransportType": "ServiceControl.Transports.Learning.LearningTransportCustomization, ServiceControl.Transports.Learning", + "MigrationEnabled": false, "ErrorLogQueue": "error.log", "ErrorQueue": "error", "ForwardErrorMessages": false, diff --git a/src/ServiceControl.UnitTests/ApprovalFiles/MigrationEngineCategorySelectionTests.Every_category_runs_in_a_fixed_order.approved.txt b/src/ServiceControl.UnitTests/ApprovalFiles/MigrationEngineCategorySelectionTests.Every_category_runs_in_a_fixed_order.approved.txt index b77ea84b2b..1e54a214bc 100644 --- a/src/ServiceControl.UnitTests/ApprovalFiles/MigrationEngineCategorySelectionTests.Every_category_runs_in_a_fixed_order.approved.txt +++ b/src/ServiceControl.UnitTests/ApprovalFiles/MigrationEngineCategorySelectionTests.Every_category_runs_in_a_fixed_order.approved.txt @@ -10,9 +10,8 @@ Required 9: LicensingThroughput, after LicensingEndpoints Required 10: LicensingReportMasks Required 11: LicensedEndpointDetails Required 12: UnresolvedAndRetryIssuedFailedMessages, with bodies +Required 14: CustomChecks +Required 15: FailedErrorImports, with bodies +Required 16: GroupComments, after UnresolvedAndRetryIssuedFailedMessages Optional 1: EventLog -Optional 2: CustomChecks -Optional 3: FailedErrorImports, with bodies -Optional 4: FailedMessageEdits -Optional 5: ArchivedAndResolvedFailedMessages, with bodies -Optional 6: GroupComments, after ArchivedAndResolvedFailedMessages \ No newline at end of file +Optional 2: ArchivedAndResolvedFailedMessages, with bodies \ No newline at end of file diff --git a/src/ServiceControl.UnitTests/ApprovalFiles/MigrationEngineCategorySelectionTests.Only_configured_optional_categories_are_selected.approved.txt b/src/ServiceControl.UnitTests/ApprovalFiles/MigrationEngineCategorySelectionTests.Only_configured_optional_categories_are_selected.approved.txt index 6c96bee11f..d0b5240bad 100644 --- a/src/ServiceControl.UnitTests/ApprovalFiles/MigrationEngineCategorySelectionTests.Only_configured_optional_categories_are_selected.approved.txt +++ b/src/ServiceControl.UnitTests/ApprovalFiles/MigrationEngineCategorySelectionTests.Only_configured_optional_categories_are_selected.approved.txt @@ -1,6 +1,4 @@ -Configured: GroupComments, EventLog, ArchivedAndResolvedFailedMessages +Configured: EventLog Runs as: Optional 1: EventLog -Optional 5: ArchivedAndResolvedFailedMessages, with bodies -Optional 6: GroupComments, after ArchivedAndResolvedFailedMessages -Never copied: CustomChecks, FailedErrorImports, FailedMessageEdits \ No newline at end of file +Never copied: ArchivedAndResolvedFailedMessages \ No newline at end of file diff --git a/src/ServiceControl.UnitTests/Migration/CheckpointMigrationStateTests.cs b/src/ServiceControl.UnitTests/Migration/CheckpointMigrationStateTests.cs new file mode 100644 index 0000000000..7a71eb1c2d --- /dev/null +++ b/src/ServiceControl.UnitTests/Migration/CheckpointMigrationStateTests.cs @@ -0,0 +1,111 @@ +#nullable enable +namespace ServiceControl.UnitTests.Migration; + +using System; +using System.Linq; +using System.Threading; +using System.Threading.Tasks; +using NUnit.Framework; +using ServiceControl.Persistence.DataMigration; +using ServiceControl.UnitTests.Migration.Fakes; + +[TestFixture] +class CheckpointMigrationStateTests +{ + [Test] + public async Task An_instance_that_has_never_migrated_has_no_store_and_nothing_incomplete() + { + var state = new CheckpointMigrationState(); + + await state.Seed(["EndpointSettings"], CancellationToken.None); + + Assert.That(state.AnyCategoryIncomplete, Is.False); + } + + [Test] + public async Task Selected_categories_that_all_finished_leave_the_guard_off() + { + var store = new InMemoryMigrationCheckpointStore(); + await store.Upsert(Checkpoint("KnownEndpoints", MigrationCategoryState.Complete)); + await store.Upsert(Checkpoint("EndpointSettings", MigrationCategoryState.Complete)); + + var state = new CheckpointMigrationState(store); + await state.Seed(["KnownEndpoints", "EndpointSettings"], CancellationToken.None); + + Assert.That(state.AnyCategoryIncomplete, Is.False); + } + + [TestCase(MigrationCategoryState.InProgress)] + [TestCase(MigrationCategoryState.NotStarted)] + [TestCase(MigrationCategoryState.Halted)] + [TestCase(MigrationCategoryState.Blocked)] + [TestCase(MigrationCategoryState.CompleteWithErrors)] + public async Task One_selected_category_short_of_finished_holds_the_guard_on(MigrationCategoryState unfinished) + { + var store = new InMemoryMigrationCheckpointStore(); + await store.Upsert(Checkpoint("KnownEndpoints", MigrationCategoryState.Complete)); + await store.Upsert(Checkpoint("EndpointSettings", unfinished)); + + var state = new CheckpointMigrationState(store); + await state.Seed(["KnownEndpoints", "EndpointSettings"], CancellationToken.None); + + Assert.That(state.AnyCategoryIncomplete, Is.True); + } + + [Test] + public async Task A_selected_category_with_no_row_at_all_has_not_started() + { + var store = new InMemoryMigrationCheckpointStore(); + await store.Upsert(Checkpoint("KnownEndpoints", MigrationCategoryState.Complete)); + + var state = new CheckpointMigrationState(store); + await state.Seed(["KnownEndpoints", "EndpointSettings"], CancellationToken.None); + + Assert.That(state.AnyCategoryIncomplete, Is.True); + } + + [Test] + public async Task An_unselected_category_that_never_ran_does_not_hold_the_guard_on_forever() + { + var store = new InMemoryMigrationCheckpointStore(); + await store.Upsert(new MigrationCheckpoint("EndpointSettings", MigrationCategoryState.Complete, null, 3, 0, 3, null, null, null, null, null)); + await store.Upsert(new MigrationCheckpoint("EventLog", MigrationCategoryState.NotStarted, null, 0, 0, null, null, null, null, null, null)); + + var state = new CheckpointMigrationState(store); + await state.Seed(["EndpointSettings"], CancellationToken.None); + + Assert.That(state.AnyCategoryIncomplete, Is.False); + } + + [Test] + public void Exactly_complete_and_abandoned_are_finished() => + Assert.That( + Enum.GetValues().Where(state => state.IsFinished()), + Is.EquivalentTo(new[] { MigrationCategoryState.Complete, MigrationCategoryState.Abandoned })); + + [Test] + public void Exactly_halted_and_complete_with_errors_are_failed() => + Assert.That( + Enum.GetValues().Where(state => state.IsFailed()), + Is.EquivalentTo(new[] { MigrationCategoryState.Halted, MigrationCategoryState.CompleteWithErrors })); + + [Test] + public void Exactly_the_two_harmless_reasons_are_harmless_and_every_other_skip_reason_is_a_fault() => + Assert.That( + Enum.GetValues().Where(reason => reason.IsBenign()), + Is.EquivalentTo(new[] { MigrationSkipReason.PastRetention, MigrationSkipReason.BlankGroupComment }), + "a harmless reason is exempt from the halt threshold and the settle rule, so any number of rows lost to one settles the category Done"); + + [Test] + public void Unknown_stays_the_last_skip_reason() => + Assert.That(Enum.GetValues().Last(), Is.EqualTo(MigrationSkipReason.Unknown)); + + [Test] + public void Exactly_the_reasons_no_retry_can_fix_are_permanent() => + Assert.That( + Enum.GetValues().Where(reason => reason.IsPermanent()), + Is.EquivalentTo(new[] { MigrationSkipReason.RequiredValueMissing })); + + static MigrationCheckpoint Checkpoint(string categoryId, MigrationCategoryState state) => + new(categoryId, state, null, 0, 0, null, null, null, null, null, null); +} diff --git a/src/ServiceControl.UnitTests/Migration/ClosedWindowProgressTests.cs b/src/ServiceControl.UnitTests/Migration/ClosedWindowProgressTests.cs new file mode 100644 index 0000000000..e381874eac --- /dev/null +++ b/src/ServiceControl.UnitTests/Migration/ClosedWindowProgressTests.cs @@ -0,0 +1,232 @@ +#nullable enable +namespace ServiceControl.UnitTests.Migration; + +using System; +using System.Linq; +using System.Threading; +using System.Threading.Tasks; +using Microsoft.Extensions.Logging; +using NUnit.Framework; +using ServiceControl.Migration; +using ServiceControl.Persistence.DataMigration; +using ServiceControl.UnitTests.Migration.Fakes; + +// No acceptance test can reach the watchdog: the host copies on the real clock and nobody waits half an hour. +[TestFixture] +class ClosedWindowProgressTests +{ + static readonly TimeSpan PollInterval = MigrationStartup.ClosedWindowProgress.PollInterval; + static readonly TimeSpan StallLimit = MigrationStartup.ClosedWindowProgress.StallLimit; + + [Test] + public async Task A_category_that_commits_nothing_for_the_stall_limit_is_stopped() + { + var clock = new TimerRecordingTimeProvider(); + var store = new PollObservingCheckpointStore(clock); + await store.Upsert(InProgress(MigrationCategoryIds.KnownEndpoints, lastProgressAt: clock.GetUtcNow().UtcDateTime)); + using var stop = new CancellationTokenSource(); + var progress = await StartWatching(store, clock, MigrationCategoryIds.KnownEndpoints, stop); + + clock.Advance(StallLimit + PollInterval); + + Assert.That(stop.Token.WaitHandle.WaitOne(TimeSpan.FromSeconds(10)), Is.True, "a category that committed nothing for longer than the limit was never stopped"); + await progress.DisposeAsync(); + } + + // A call that ignores its token never returns, so the row stays in progress and every later poll sees the same stall. + [Test] + public async Task A_category_already_stopped_is_reported_once_however_long_it_stays_stuck() + { + var clock = new TimerRecordingTimeProvider(); + var store = new PollObservingCheckpointStore(clock); + await store.Upsert(InProgress(MigrationCategoryIds.KnownEndpoints, lastProgressAt: clock.GetUtcNow().UtcDateTime)); + var logger = new CapturingLogger(); + using var stop = new CancellationTokenSource(); + var progress = await StartWatching(store, clock, MigrationCategoryIds.KnownEndpoints, stop, logger: logger); + + clock.Advance(StallLimit + PollInterval); + await store.WaitForAPollAtTheCurrentTime(); + clock.Advance(PollInterval); + await store.WaitForAPollAtTheCurrentTime(); + await progress.DisposeAsync(); + + Assert.That(logger.Entries.Count(entry => entry.Level == LogLevel.Error), Is.EqualTo(1), "the same stall was reported again on the next poll"); + } + + [Test] + public async Task A_resumed_row_carrying_the_previous_runs_stamp_is_not_a_stall() + { + // The row a killed run left behind keeps its last stamp, so an operator restarting hours later + // would otherwise have a healthy copy cancelled on the first tick, every time. + var clock = new TimerRecordingTimeProvider(); + var store = new PollObservingCheckpointStore(clock); + await store.Upsert(InProgress(MigrationCategoryIds.KnownEndpoints, lastProgressAt: clock.GetUtcNow().UtcDateTime - TimeSpan.FromHours(2))); + using var stop = new CancellationTokenSource(); + var progress = await StartWatching(store, clock, MigrationCategoryIds.KnownEndpoints, stop); + + clock.Advance(PollInterval); + await store.WaitForAPollAtTheCurrentTime(); + await progress.DisposeAsync(); + + Assert.That(stop.IsCancellationRequested, Is.False, "the window runs from when this category's run began, not from a stamp the previous run left"); + } + + [Test] + public async Task A_row_other_than_the_running_category_is_not_judged() + { + // Both were attempted, but only KnownEndpoints is running, and it has no row yet. + var clock = new TimerRecordingTimeProvider(); + var store = new PollObservingCheckpointStore(clock); + await store.Upsert(InProgress(MigrationCategoryIds.EndpointSettings, lastProgressAt: clock.GetUtcNow().UtcDateTime)); + using var stop = new CancellationTokenSource(); + var progress = await StartWatching(store, clock, MigrationCategoryIds.KnownEndpoints, stop, [MigrationCategoryIds.KnownEndpoints, MigrationCategoryIds.EndpointSettings]); + + clock.Advance(StallLimit + PollInterval); + await store.WaitForAPollAtTheCurrentTime(); + await progress.DisposeAsync(); + + Assert.That(stop.IsCancellationRequested, Is.False, "the running category was stopped over a row nobody is copying"); + } + + [Test] + public async Task Committing_nothing_for_exactly_the_stall_limit_is_not_a_stall() + { + var clock = new TimerRecordingTimeProvider(); + var store = new PollObservingCheckpointStore(clock); + await store.Upsert(InProgress(MigrationCategoryIds.KnownEndpoints, lastProgressAt: clock.GetUtcNow().UtcDateTime)); + using var stop = new CancellationTokenSource(); + var progress = await StartWatching(store, clock, MigrationCategoryIds.KnownEndpoints, stop); + + clock.Advance(StallLimit); + await store.WaitForAPollAtTheCurrentTime(); + await progress.DisposeAsync(); + + Assert.That(stop.IsCancellationRequested, Is.False, "the limit is the point at which a copy has not yet stalled"); + } + + [Test] + public async Task A_category_that_has_committed_nothing_since_its_run_began_is_stopped() + { + // The engine stamps a category when it marks it running, so the window covers the first read, which is + // where a copy that never gets going actually hangs. When the category first started does not matter. + var clock = new TimerRecordingTimeProvider(); + var store = new PollObservingCheckpointStore(clock); + await store.Upsert(InProgress(MigrationCategoryIds.KnownEndpoints, lastProgressAt: clock.GetUtcNow().UtcDateTime, startedAt: clock.GetUtcNow().UtcDateTime - TimeSpan.FromDays(3))); + using var stop = new CancellationTokenSource(); + var progress = await StartWatching(store, clock, MigrationCategoryIds.KnownEndpoints, stop); + + clock.Advance(StallLimit + PollInterval); + + Assert.That(stop.Token.WaitHandle.WaitOne(TimeSpan.FromSeconds(10)), Is.True, "a category that has committed nothing since its run began was never stopped"); + await progress.DisposeAsync(); + } + + [Test] + public async Task A_row_an_older_build_left_without_a_stamp_is_still_watched() + { + var clock = new TimerRecordingTimeProvider(); + var store = new PollObservingCheckpointStore(clock); + await store.Upsert(InProgress(MigrationCategoryIds.KnownEndpoints, lastProgressAt: null, startedAt: clock.GetUtcNow().UtcDateTime - TimeSpan.FromDays(3))); + using var stop = new CancellationTokenSource(); + var progress = await StartWatching(store, clock, MigrationCategoryIds.KnownEndpoints, stop); + + clock.Advance(StallLimit + PollInterval); + + Assert.That(stop.Token.WaitHandle.WaitOne(TimeSpan.FromSeconds(10)), Is.True, "an unstamped row was left unwatched for ever"); + await progress.DisposeAsync(); + } + + [Test] + public async Task A_category_this_run_finished_is_not_stopped_for_having_gone_quiet() + { + // The watch moves to the next category only when that one starts, so a settled row can still be the one watched. + var clock = new TimerRecordingTimeProvider(); + var store = new PollObservingCheckpointStore(clock); + await store.Upsert(Settled(MigrationCategoryIds.KnownEndpoints, MigrationCategoryState.Complete, lastProgressAt: clock.GetUtcNow().UtcDateTime)); + using var stop = new CancellationTokenSource(); + var progress = await StartWatching(store, clock, MigrationCategoryIds.KnownEndpoints, stop); + + clock.Advance(StallLimit + PollInterval); + await store.WaitForAPollAtTheCurrentTime(); + await progress.DisposeAsync(); + + Assert.That(stop.IsCancellationRequested, Is.False, "a category that finished was stopped"); + } + + [Test] + public async Task A_halted_category_this_run_attempted_is_not_stopped_for_having_gone_quiet() + { + // A halt settles the row and leaves its stamp behind. The gate is what refuses the host over it, not the watchdog. + var clock = new TimerRecordingTimeProvider(); + var store = new PollObservingCheckpointStore(clock); + await store.Upsert(Settled(MigrationCategoryIds.KnownEndpoints, MigrationCategoryState.Halted, lastProgressAt: clock.GetUtcNow().UtcDateTime)); + using var stop = new CancellationTokenSource(); + var progress = await StartWatching(store, clock, MigrationCategoryIds.KnownEndpoints, stop); + + clock.Advance(StallLimit + PollInterval); + await store.WaitForAPollAtTheCurrentTime(); + await progress.DisposeAsync(); + + Assert.That(stop.IsCancellationRequested, Is.False, "a halted category was stopped"); + } + + // The watch has no total limit, so a copy that keeps committing outlives any length of run. + [Test] + public async Task A_category_still_committing_batches_is_never_stopped_however_long_it_takes() + { + var clock = new TimerRecordingTimeProvider(); + var store = new PollObservingCheckpointStore(clock); + var committed = await store.Upsert(InProgress(MigrationCategoryIds.KnownEndpoints, lastProgressAt: clock.GetUtcNow().UtcDateTime)); + using var stop = new CancellationTokenSource(); + var progress = await StartWatching(store, clock, MigrationCategoryIds.KnownEndpoints, stop); + + // Four times the limit, committing a batch every poll, which a total timeout would have killed long ago. + for (var elapsed = TimeSpan.Zero; elapsed < StallLimit * 4; elapsed += PollInterval) + { + clock.Advance(PollInterval); + await store.WaitForAPollAtTheCurrentTime(); + committed = await store.Upsert(committed with { LastProgressAt = clock.GetUtcNow().UtcDateTime }); + } + + await progress.DisposeAsync(); + + Assert.That(stop.IsCancellationRequested, Is.False, "a copy committing a batch every poll was stopped, so the watch is a deadline rather than a stall detector"); + } + + [Test] + public async Task A_poll_that_fails_leaves_the_watch_running_and_says_so() + { + var clock = new TimerRecordingTimeProvider(); + var store = new PollObservingCheckpointStore(clock); + var storeFailure = new InvalidOperationException("the checkpoint store is unreachable"); + store.FailReadAll(1, storeFailure); + await store.Upsert(InProgress(MigrationCategoryIds.KnownEndpoints, lastProgressAt: clock.GetUtcNow().UtcDateTime)); + var logger = new CapturingLogger(); + using var stop = new CancellationTokenSource(); + var progress = await StartWatching(store, clock, MigrationCategoryIds.KnownEndpoints, stop, logger: logger); + + clock.Advance(PollInterval); + await store.WaitForAPollAtTheCurrentTime(); + clock.Advance(StallLimit + PollInterval); + + Assert.That(stop.Token.WaitHandle.WaitOne(TimeSpan.FromSeconds(10)), Is.True, "a watch that stopped at the first blip never noticed the stall that followed"); + await progress.DisposeAsync(); + Assert.That(logger.Entries.Where(entry => entry.Level == LogLevel.Error).Select(entry => entry.Exception), Has.Member(storeFailure), "a watch that has gone deaf has to say so"); + } + + static async Task StartWatching(PollObservingCheckpointStore store, TimerRecordingTimeProvider clock, string runningCategoryId, CancellationTokenSource stop, string[]? attemptedCategoryIds = null, ILogger? logger = null) + { + var progress = new MigrationStartup.ClosedWindowProgress(store, clock, logger ?? new CapturingLogger(), attemptedCategoryIds ?? [runningCategoryId]); + progress.Watch(runningCategoryId, clock.GetUtcNow().UtcDateTime, stop); + + Assert.That(await clock.TimerCreated.WaitAsync(TimeSpan.FromSeconds(10)), Is.True, "the watch never started its timer, so advancing the clock would tick nothing"); + + return progress; + } + + static MigrationCheckpoint InProgress(string categoryId, DateTime? lastProgressAt, DateTime? startedAt = null) => + new(categoryId, MigrationCategoryState.InProgress, null, 0, 0, null, null, startedAt ?? lastProgressAt, lastProgressAt, null, null); + + static MigrationCheckpoint Settled(string categoryId, MigrationCategoryState state, DateTime lastProgressAt) => + new(categoryId, state, null, 0, 0, null, null, lastProgressAt, lastProgressAt, lastProgressAt, null); +} diff --git a/src/ServiceControl.UnitTests/Migration/CopyableCategoryIdsTests.cs b/src/ServiceControl.UnitTests/Migration/CopyableCategoryIdsTests.cs new file mode 100644 index 0000000000..c8774a9e9a --- /dev/null +++ b/src/ServiceControl.UnitTests/Migration/CopyableCategoryIdsTests.cs @@ -0,0 +1,19 @@ +namespace ServiceControl.UnitTests.Migration; + +using NUnit.Framework; +using ServiceControl.Migration; +using ServiceControl.Persistence.DataMigration; + +[TestFixture] +class CopyableCategoryIdsTests +{ + [Test] + public void A_category_only_one_side_supports_counts_as_not_copyable() + { + var source = new[] { MigrationCategoryIds.KnownEndpoints, MigrationCategoryIds.EndpointSettings }; + var target = new[] { MigrationCategoryIds.KnownEndpoints }; + + Assert.That(MigrationStartup.CopyableCategoryIds(source, target), Is.EquivalentTo(new[] { MigrationCategoryIds.KnownEndpoints }), + "a category with a reader and no writer would otherwise be attempted and fail partway through a customer's copy"); + } +} diff --git a/src/ServiceControl.UnitTests/Migration/Fakes/InMemoryMigrationSource.cs b/src/ServiceControl.UnitTests/Migration/Fakes/InMemoryMigrationSource.cs index 66341cbf20..ebc2dd8253 100644 --- a/src/ServiceControl.UnitTests/Migration/Fakes/InMemoryMigrationSource.cs +++ b/src/ServiceControl.UnitTests/Migration/Fakes/InMemoryMigrationSource.cs @@ -13,6 +13,7 @@ public sealed class InMemoryMigrationSource : IMigrationSource { readonly Dictionary> rowsByCategory = []; readonly Dictionary bodyFailures = []; + readonly Dictionary bodyFailuresInTurn = []; readonly Dictionary bodyReadAttempts = []; readonly Dictionary bodies = []; @@ -24,13 +25,22 @@ public sealed class InMemoryMigrationSource : IMigrationSource public void FailBodyReads(string sourceId, int times, Exception failure) => bodyFailures[sourceId] = (times, failure); - /// Makes the body read for this row behave like a host shutting down: the token is cancelled and the read throws. + /// + /// Fails the body reads for this row with one exception per attempt, in order, so a test can tell one attempt's error from another's. An attempt past the end of the list reads the body normally. + /// + public void FailBodyReadsInTurn(string sourceId, params Exception[] failuresInAttemptOrder) => bodyFailuresInTurn[sourceId] = failuresInAttemptOrder; + + /// + /// Makes the body read for this row behave like a host shutting down: the token is cancelled and the read throws. + /// public (string SourceId, CancellationTokenSource Source)? StopOnBodyRead { get; set; } public int BodyReadAttempts(string sourceId) => bodyReadAttempts.GetValueOrDefault(sourceId); public Task Open(CancellationToken cancellationToken = default) => Task.CompletedTask; + public IReadOnlyList ContributedChecks() => []; + public Task Describe(CancellationToken cancellationToken = default) => Task.FromResult(Description); public Task> Inventory(CancellationToken cancellationToken = default) => @@ -38,7 +48,7 @@ public Task> Inventory(Cancellation [.. rowsByCategory.Select(pair => new MigrationSourceInventoryEntry("memory", pair.Key, pair.Value.Count))]); public Task Count(MigrationCategory category, CancellationToken cancellationToken = default) => - Task.FromResult((long)(rowsByCategory.TryGetValue(category.Id, out var rows) ? rows.Count : 0)); + Task.FromResult(rowsByCategory.TryGetValue(category.Id, out var rows) ? rows.Count : 0); public async IAsyncEnumerable Read( MigrationCategory category, @@ -85,8 +95,19 @@ public async IAsyncEnumerable Read( throw failures.Failure; } + if (bodyFailuresInTurn.TryGetValue(sourceId, out var failuresInTurn) && attempt <= failuresInTurn.Length) + { + throw failuresInTurn[attempt - 1]; + } + return bodies.GetValueOrDefault(sourceId); } + // Every registered category, not just the seeded ones: Seed writes rowsByCategory after construction, + // and the interface promises an answer that does not change across Open. + public IReadOnlyCollection SupportedCategoryIds => [.. MigrationCategoryRegistry.All.Select(category => category.Id)]; + + public IReadOnlyDictionary DocumentTypes => MigrationCategoryRegistry.All.ToDictionary(category => category.Id, _ => typeof(object)); + public ValueTask DisposeAsync() => ValueTask.CompletedTask; } diff --git a/src/ServiceControl.UnitTests/Migration/Fakes/InMemoryMigrationTarget.cs b/src/ServiceControl.UnitTests/Migration/Fakes/InMemoryMigrationTarget.cs index cdfaff9777..84c171e912 100644 --- a/src/ServiceControl.UnitTests/Migration/Fakes/InMemoryMigrationTarget.cs +++ b/src/ServiceControl.UnitTests/Migration/Fakes/InMemoryMigrationTarget.cs @@ -2,6 +2,7 @@ namespace ServiceControl.UnitTests.Migration.Fakes; using System; using System.Collections.Generic; +using System.Linq; using System.Threading; using System.Threading.Tasks; using ServiceControl.Persistence.DataMigration; @@ -10,30 +11,60 @@ public sealed class InMemoryMigrationTarget(IMigrationCheckpointStore checkpoint { readonly Dictionary> writtenKeysByCategory = []; readonly Dictionary> writtenRowsByCategory = []; + readonly Dictionary> rowsHandedToWriteByCategory = []; readonly HashSet preExistingKeys = []; - readonly Dictionary rejectedKeys = []; + readonly Dictionary rejectedKeys = []; public int DefaultBatchSize { get; set; } = 3; public string NoBatchSizeFor { get; set; } public int? FailOnCallNumber { get; set; } - /// What FailOnCallNumber throws, when the default simulated failure is the wrong shape for the test. + /// + /// What FailOnCallNumber throws, when the default simulated failure is the wrong shape for the test. + /// public Exception FailWith { get; set; } - /// Cancels the token on this call and then writes normally, so the stop surfaces from the source's next batch. + /// + /// How many rows each write leaves out of the outcome it reports and commits, without failing, so the two still agree with each other. + /// + public int UnderReportBy { get; set; } + + /// + /// Runs at the start of every write, before any simulated failure, with the checkpoint the engine is extending, so a test can watch a stamp move between batches. + /// + public Action BeforeWrite { get; set; } + + /// + /// Cancels the token on this call and then writes normally, so the stop surfaces from the source's next batch. + /// public (int CallNumber, CancellationTokenSource Source)? CancelOnCall { get; set; } + /// + /// Cancels the token on this call and throws instead of writing, so the stop surfaces from the write itself. + /// public (int CallNumber, CancellationTokenSource Source)? StopOnCall { get; set; } + /// + /// Waits on this call until the token is cancelled and then throws, as a write stuck on a database that stopped answering does. + /// + public int? HangOnCall { get; set; } int callCount; public void SeedExistingKey(string sourceId) => preExistingKeys.Add(sourceId); - public void RejectKey(string sourceId, MigrationSkipReason reason, bool benign = false) => rejectedKeys[sourceId] = (reason, benign); + public void RejectKey(string sourceId, MigrationSkipReason reason) => rejectedKeys[sourceId] = reason; public IReadOnlyList WrittenRows(string categoryId) => writtenRowsByCategory.TryGetValue(categoryId, out var rows) ? rows : []; - public int BatchSizeFor(MigrationCategory category) => - category.Id == NoBatchSizeFor ? throw new InvalidOperationException($"No batch size is mapped for category {category.Id}") : DefaultBatchSize; + /// + /// Every row the engine sent to a write that ran, in the order it sent them, including the rows this fake then skipped or found already present. It is the only way to see the same row sent twice, because WrittenRows de-duplicates as the real targets do. Empty for a category never written to. + /// + public IReadOnlyList RowsHandedToWrite(string categoryId) => + rowsHandedToWriteByCategory.TryGetValue(categoryId, out var rows) ? rows : []; + + public Task Open(CancellationToken cancellationToken = default) => Task.CompletedTask; + + public Task BatchSizeFor(MigrationCategory category, CancellationToken cancellationToken = default) => + category.Id == NoBatchSizeFor ? throw new InvalidOperationException($"No batch size is mapped for category {category.Id}") : Task.FromResult(DefaultBatchSize); public async Task Write( MigrationCategory category, @@ -42,6 +73,12 @@ public async Task Write( CancellationToken cancellationToken = default) { callCount++; + BeforeWrite?.Invoke(checkpointToExtend); + + if (HangOnCall == callCount) + { + await Task.Delay(Timeout.Infinite, cancellationToken); + } if (StopOnCall is { } stop && stop.CallNumber == callCount) { @@ -62,22 +99,22 @@ public async Task Write( var keys = writtenKeysByCategory.TryGetValue(category.Id, out var existingKeys) ? existingKeys : writtenKeysByCategory[category.Id] = []; var rows = writtenRowsByCategory.TryGetValue(category.Id, out var existingRows) ? existingRows : writtenRowsByCategory[category.Id] = []; + // Recorded here rather than at the top of the method: a write that threw above never committed, so + // the restart that sends its batch again is doing the right thing. + var handedOver = rowsHandedToWriteByCategory.TryGetValue(category.Id, out var existingHandedOver) ? existingHandedOver : rowsHandedToWriteByCategory[category.Id] = []; + handedOver.AddRange(batch.Rows); + var copied = 0; var alreadyPresent = 0; - var benignSkipped = 0; var skippedIds = new List(); var skipReasons = new Dictionary(); foreach (var row in batch.Rows) { - if (rejectedKeys.TryGetValue(row.SourceId, out var rejection)) + if (rejectedKeys.TryGetValue(row.SourceId, out var reason)) { skippedIds.Add(row.SourceId); - skipReasons[rejection.Reason] = skipReasons.GetValueOrDefault(rejection.Reason) + 1; - if (rejection.Benign) - { - benignSkipped++; - } + skipReasons[reason] = skipReasons.GetValueOrDefault(reason) + 1; continue; } @@ -91,15 +128,21 @@ public async Task Write( copied++; } - // Extended and saved in the same operation as the rows, as the real targets do, so what lands - // is this batch's real split rather than a provisional one the next save has to correct. + copied -= UnderReportBy; + + // The real targets extend and save the checkpoint in the transaction that writes the rows, + // so this fake saves it here too. var saved = await checkpointStore.Upsert( checkpointToExtend.Extend(copied, skippedIds.Count, alreadyPresent, skipReasons), cancellationToken); - return new MigrationWriteResult(saved, copied, skippedIds.Count, skippedIds, alreadyPresent, skipReasons, benignSkipped); + return new MigrationWriteResult(saved, copied, skippedIds.Count, skippedIds, alreadyPresent, skipReasons); } public Task Count(MigrationCategory category, CancellationToken cancellationToken = default) => Task.FromResult((long)(writtenRowsByCategory.TryGetValue(category.Id, out var rows) ? rows.Count : 0)); + + public IReadOnlyCollection SupportedCategoryIds => [.. MigrationCategoryRegistry.All.Select(category => category.Id)]; + + public IReadOnlyDictionary DocumentTypes => MigrationCategoryRegistry.All.ToDictionary(category => category.Id, _ => typeof(object)); } diff --git a/src/ServiceControl.UnitTests/Migration/Fakes/PollObservingCheckpointStore.cs b/src/ServiceControl.UnitTests/Migration/Fakes/PollObservingCheckpointStore.cs new file mode 100644 index 0000000000..147fac3a8f --- /dev/null +++ b/src/ServiceControl.UnitTests/Migration/Fakes/PollObservingCheckpointStore.cs @@ -0,0 +1,64 @@ +#nullable enable +namespace ServiceControl.UnitTests.Migration.Fakes; + +using System; +using System.Collections.Concurrent; +using System.Collections.Generic; +using System.Threading; +using System.Threading.Tasks; +using NUnit.Framework; +using ServiceControl.Persistence.DataMigration; + +// Records when each poll read the rows, so a test can wait for a poll that saw the state it is judging. +// Only the stall watchdog calls ReadAll, so every ReadAll is one of its polls. +public sealed class PollObservingCheckpointStore(TimeProvider clock) : IMigrationCheckpointStore +{ + readonly InMemoryMigrationCheckpointStore inner = new(); + readonly SemaphoreSlim polled = new(0); + readonly ConcurrentQueue polledAt = new(); + Exception? readAllFailure; + int readAllFailuresLeft; + + public void FailReadAll(int times, Exception failure) + { + readAllFailuresLeft = times; + readAllFailure = failure; + } + + /// + /// Waits until a poll has read the rows at the clock's current time, and fails the test if none does within ten + /// seconds. Polls at earlier times are passed over. + /// + public async Task WaitForAPollAtTheCurrentTime() + { + var now = clock.GetUtcNow().UtcDateTime; + + while (true) + { + Assert.That(await polled.WaitAsync(TimeSpan.FromSeconds(10)), Is.True, $"no poll read the checkpoints at {now:O}; a watch that stopped polling never reaches one"); + + if (polledAt.TryDequeue(out var at) && at == now) + { + return; + } + } + } + + public Task> ReadAll(CancellationToken cancellationToken = default) + { + polledAt.Enqueue(clock.GetUtcNow().UtcDateTime); + polled.Release(); + + if (readAllFailuresLeft > 0) + { + readAllFailuresLeft--; + return Task.FromException>(readAllFailure!); + } + + return inner.ReadAll(cancellationToken); + } + + public Task Read(string categoryId, CancellationToken cancellationToken = default) => inner.Read(categoryId, cancellationToken); + + public Task Upsert(MigrationCheckpoint checkpoint, CancellationToken cancellationToken = default) => inner.Upsert(checkpoint, cancellationToken); +} diff --git a/src/ServiceControl.UnitTests/Migration/HaltThresholdTests.cs b/src/ServiceControl.UnitTests/Migration/HaltThresholdTests.cs index fbe2bebe97..ae309d6c8b 100644 --- a/src/ServiceControl.UnitTests/Migration/HaltThresholdTests.cs +++ b/src/ServiceControl.UnitTests/Migration/HaltThresholdTests.cs @@ -61,6 +61,16 @@ public void A_hair_past_the_proportion_halts_once_the_floor_is_behind_it() Assert.That(exceeded, Is.True); } + [Test] + public void A_large_category_losing_a_small_share_does_not_halt_once_past_the_floor() + { + // 101 of 5,000,000 is 0.002%. The floor is already behind it, so only the proportion can stop a halt + // here, and a rule that halted on the floor alone would stop a healthy copy of a huge category. + var exceeded = HaltThreshold.Exceeded(skippedCount: 101, totalCount: 5_000_000, percentThreshold: 5, minimumFloor: 100); + + Assert.That(exceeded, Is.False); + } + [Test] public void A_skip_count_with_nothing_processed_never_halts_and_never_divides_by_zero() { diff --git a/src/ServiceControl.UnitTests/Migration/MigrationCategoryRegistryTests.cs b/src/ServiceControl.UnitTests/Migration/MigrationCategoryRegistryTests.cs index 9b5e637ec3..c82e610f09 100644 --- a/src/ServiceControl.UnitTests/Migration/MigrationCategoryRegistryTests.cs +++ b/src/ServiceControl.UnitTests/Migration/MigrationCategoryRegistryTests.cs @@ -12,12 +12,12 @@ namespace ServiceControl.UnitTests.Migration; class MigrationCategoryRegistryTests { [Test] - public void Contains_all_eighteen_categories_with_unique_ids() + public void Contains_all_seventeen_categories_with_unique_ids() { using (Assert.EnterMultipleScope()) { - Assert.That(MigrationCategoryRegistry.All, Has.Count.EqualTo(18)); - Assert.That(MigrationCategoryRegistry.All.Select(c => c.Id).Distinct().Count(), Is.EqualTo(18)); + Assert.That(MigrationCategoryRegistry.All, Has.Count.EqualTo(17)); + Assert.That(MigrationCategoryRegistry.All.Select(c => c.Id).Distinct().Count(), Is.EqualTo(17)); } } @@ -35,7 +35,8 @@ public void Required_categories_are_ordered_exactly_as_the_contract_lists_them() "KnownEndpoints", "EndpointSettings", "MessageRedirects", "Subscriptions", "NotificationSettings", "TrialEndDate", "RetryOperations", "LicensingEndpoints", "LicensingThroughput", "LicensingReportMasks", - "LicensedEndpointDetails", "UnresolvedAndRetryIssuedFailedMessages" + "LicensedEndpointDetails", "UnresolvedAndRetryIssuedFailedMessages", + "CustomChecks", "FailedErrorImports", "GroupComments" })); } @@ -50,8 +51,7 @@ public void Optional_categories_are_ordered_exactly_as_the_contract_lists_them() Assert.That(optionalIds, Is.EqualTo(new[] { - "EventLog", "CustomChecks", "FailedErrorImports", - "FailedMessageEdits", "ArchivedAndResolvedFailedMessages", "GroupComments" + "EventLog", "ArchivedAndResolvedFailedMessages" })); } @@ -71,7 +71,7 @@ public void Three_categories_declare_a_MustFollow() ["EndpointSettings"] = "KnownEndpoints", // A group comment whose failed messages have not arrived reads as an orphan to the // retention sweeper, which deletes it. - ["GroupComments"] = "ArchivedAndResolvedFailedMessages" + ["GroupComments"] = "UnresolvedAndRetryIssuedFailedMessages" })); } diff --git a/src/ServiceControl.UnitTests/Migration/MigrationContractShapeTests.cs b/src/ServiceControl.UnitTests/Migration/MigrationContractShapeTests.cs index 3df3d4e123..8b981b0665 100644 --- a/src/ServiceControl.UnitTests/Migration/MigrationContractShapeTests.cs +++ b/src/ServiceControl.UnitTests/Migration/MigrationContractShapeTests.cs @@ -61,4 +61,23 @@ public void MigrationCheckpoint_AlreadyPresentCount_defaults_to_zero() Assert.That(checkpoint.AlreadyPresentCount, Is.Zero); } + + [Test] + public void MigrationCheckpoint_StartedWindowSeconds_defaults_to_null() + { + var checkpoint = new MigrationCheckpoint( + CategoryId: "EndpointSettings", + State: MigrationCategoryState.NotStarted, + Cursor: null, + CopiedCount: 0, + SkippedCount: 0, + SourceTotal: null, + SkipReasons: null, + StartedAt: null, + LastProgressAt: null, + SettledAt: null, + LastError: null); + + Assert.That(checkpoint.StartedWindowSeconds, Is.Null); + } } diff --git a/src/ServiceControl.UnitTests/Migration/MigrationEnabledSettingsTests.cs b/src/ServiceControl.UnitTests/Migration/MigrationEnabledSettingsTests.cs new file mode 100644 index 0000000000..4d82cac0f8 --- /dev/null +++ b/src/ServiceControl.UnitTests/Migration/MigrationEnabledSettingsTests.cs @@ -0,0 +1,52 @@ +namespace ServiceControl.UnitTests.Migration; + +using System; +using System.Linq; +using System.Reflection; +using NUnit.Framework; +using ServiceBus.Management.Infrastructure.Settings; +using ServiceControl.Configuration; +using ServiceControl.Persistence.DataMigration; + +[TestFixture] +[NonParallelizable] +class MigrationEnabledSettingsTests +{ + static Settings NewSettings() => + new(transportType: "LearningTransport", persisterType: "RavenDB", errorRetentionPeriod: TimeSpan.FromDays(10)); + + [TearDown] + public void ClearEnvironmentVariables() + { + Environment.SetEnvironmentVariable("SERVICECONTROL_MIGRATION_ENABLED", null); + } + + [Test] + public void Migration_defaults_to_off() + { + Assert.That(NewSettings().MigrationEnabled, Is.False); + } + + [Test] + public void MigrationEnabled_is_read_from_the_environment() + { + Environment.SetEnvironmentVariable("SERVICECONTROL_MIGRATION_ENABLED", "true"); + + Assert.That(NewSettings().MigrationEnabled, Is.True); + } + + [Test] + public void Only_MigrationEngineOptions_parses_the_engine_settings_keys() + { + var parsers = typeof(Settings).Assembly.GetTypes() + .Concat(typeof(MigrationEngineOptions).Assembly.GetTypes()) + .Where(type => type.GetMethods(BindingFlags.Public | BindingFlags.NonPublic | BindingFlags.Static) + .Any(method => method.Name is "FromSettings" or "Read" or "Resolve" + && method.GetParameters() is [{ ParameterType.Name: nameof(SettingsRootNamespace) }, ..])) + .Select(type => type.Name) + .ToArray(); + + Assert.That(parsers, Is.EquivalentTo(new[] { nameof(MigrationEngineOptions) }), + "Migration/ThrottlePauseMilliseconds and its neighbours have exactly one parser. A second one drifts its defaults and its refusal message away from this one, and nothing fails until a customer types a window wrong."); + } +} diff --git a/src/ServiceControl.UnitTests/Migration/MigrationEngineBodyRetryTests.cs b/src/ServiceControl.UnitTests/Migration/MigrationEngineBodyRetryTests.cs index dc938d6543..dbc94ad95a 100644 --- a/src/ServiceControl.UnitTests/Migration/MigrationEngineBodyRetryTests.cs +++ b/src/ServiceControl.UnitTests/Migration/MigrationEngineBodyRetryTests.cs @@ -114,10 +114,11 @@ public async Task The_configured_backoff_is_waited_between_body_read_attempts() // millisecond and the message is skipped for an outage it would have survived. var category = MigrationCategoryRegistry.Find("UnresolvedAndRetryIssuedFailedMessages")!; var source = new InMemoryMigrationSource(); - source.Seed(category.Id, Row("msg-1")); - var body = new MigrationBody(new byte[] { 1 }, "text/plain"); - source.SetBody("msg-1", body); - source.FailBodyReads("msg-1", MigrationEngine.MaxBodyReadAttempts - 1, new TimeoutException("body store unreachable")); + source.Seed(category.Id, Row("msg-1"), Row("msg-2")); + // Every attempt on msg-1 fails, so the guard against a wait after the final one is the only thing + // keeping the third timer from being created. + source.FailBodyReads("msg-1", MigrationEngine.MaxBodyReadAttempts, new TimeoutException("body store unreachable")); + source.SetBody("msg-2", new MigrationBody(new byte[] { 2 }, "text/plain")); var checkpointStore = new InMemoryMigrationCheckpointStore(); var target = new InMemoryMigrationTarget(checkpointStore); var clock = new TimerRecordingTimeProvider(); @@ -127,7 +128,8 @@ public async Task The_configured_backoff_is_waited_between_body_read_attempts() var runTask = engine.RunCategoryAsync(category); - // Two failures, so a wait after each before the attempt that succeeds. + // A wait after the first two failures only. A third would never complete, because the clock is + // not advanced again and the run below would time out waiting for it. for (var waitNumber = 1; waitNumber <= MigrationEngine.MaxBodyReadAttempts - 1; waitNumber++) { Assert.That(await clock.TimerCreated.WaitAsync(TimeSpan.FromSeconds(5)), Is.True, $"backoff {waitNumber} never started"); @@ -139,20 +141,24 @@ public async Task The_configured_backoff_is_waited_between_body_read_attempts() using (Assert.EnterMultipleScope()) { - Assert.That(checkpoint.State, Is.EqualTo(MigrationCategoryState.Complete)); + Assert.That(checkpoint.State, Is.EqualTo(MigrationCategoryState.CompleteWithErrors)); + Assert.That(checkpoint.SkippedCount, Is.EqualTo(1), "msg-1 was given up on, so all three attempts ran"); Assert.That(clock.DueTimes, Is.EqualTo(new[] { backoff, backoff }), "no wait after the final attempt, which has nothing left to retry"); - Assert.That(target.WrittenRows(category.Id).Single().Body, Is.EqualTo(body)); } } [Test] public async Task The_skip_warning_for_an_unreadable_body_carries_the_last_attempts_exception() { + // The first failure of a retry run is usually a transient connect error. The last one says what the + // body store was doing when the message was given up on, and it is all the operator gets. var category = MigrationCategoryRegistry.Find("UnresolvedAndRetryIssuedFailedMessages")!; var source = new InMemoryMigrationSource(); source.Seed(category.Id, Row("msg-1")); - var failure = new TimeoutException("body store unreachable"); - source.FailBodyReads("msg-1", MigrationEngine.MaxBodyReadAttempts, failure); + var failures = Enumerable.Range(1, MigrationEngine.MaxBodyReadAttempts) + .Select(attempt => new TimeoutException($"body store unreachable on attempt {attempt}")) + .ToArray(); + source.FailBodyReadsInTurn("msg-1", failures); var checkpointStore = new InMemoryMigrationCheckpointStore(); var target = new InMemoryMigrationTarget(checkpointStore); var options = new MigrationEngineOptions(TimeSpan.Zero, 5, 100, []) { BodyRetryBackoff = TimeSpan.Zero }; @@ -161,7 +167,7 @@ public async Task The_skip_warning_for_an_unreadable_body_carries_the_last_attem await engine.RunCategoryAsync(category); - Assert.That(logger.Entries.Single(e => e.Message.StartsWith("Skipped msg-1")).Exception, Is.SameAs(failure)); + Assert.That(logger.Entries.Single(e => e.Message.StartsWith("Skipped msg-1")).Exception, Is.SameAs(failures[^1])); } [Test] diff --git a/src/ServiceControl.UnitTests/Migration/MigrationEngineCategorySelectionTests.cs b/src/ServiceControl.UnitTests/Migration/MigrationEngineCategorySelectionTests.cs index 9c74a7b605..15cf13aaad 100644 --- a/src/ServiceControl.UnitTests/Migration/MigrationEngineCategorySelectionTests.cs +++ b/src/ServiceControl.UnitTests/Migration/MigrationEngineCategorySelectionTests.cs @@ -3,6 +3,7 @@ namespace ServiceControl.UnitTests.Migration; using System; using System.Collections.Generic; using System.Linq; +using System.Threading.Tasks; using Microsoft.Extensions.Logging.Abstractions; using Microsoft.Extensions.Time.Testing; using NUnit.Framework; @@ -13,9 +14,9 @@ namespace ServiceControl.UnitTests.Migration; [TestFixture] class MigrationEngineCategorySelectionTests { - static MigrationEngine BuildEngine(IReadOnlyCollection selectedOptionalIds, out InMemoryMigrationCheckpointStore checkpointStore) + static MigrationEngine BuildEngine(IReadOnlyCollection selectedOptionalIds) { - checkpointStore = new InMemoryMigrationCheckpointStore(); + var checkpointStore = new InMemoryMigrationCheckpointStore(); var target = new InMemoryMigrationTarget(checkpointStore); var options = new MigrationEngineOptions(TimeSpan.FromSeconds(1), 5, 100, selectedOptionalIds); return new MigrationEngine(new InMemoryMigrationSource(), target, checkpointStore, new FakeTimeProvider(), options, NullLogger.Instance); @@ -33,7 +34,7 @@ public void Every_category_runs_in_a_fixed_order() .Where(category => category.Kind == MigrationCategoryKind.Optional) .Select(category => category.Id) .ToArray(); - var engine = BuildEngine(everyOptionalId, out _); + var engine = BuildEngine(everyOptionalId); var runOrder = engine.SelectCategories(MigrationCategoryKind.Required) .Concat(engine.SelectCategories(MigrationCategoryKind.Optional)) @@ -45,8 +46,8 @@ public void Every_category_runs_in_a_fixed_order() [Test] public void Only_configured_optional_categories_are_selected() { - string[] configured = [MigrationCategoryIds.GroupComments, MigrationCategoryIds.EventLog, MigrationCategoryIds.ArchivedAndResolvedFailedMessages]; - var engine = BuildEngine(configured, out _); + string[] configured = [MigrationCategoryIds.EventLog]; + var engine = BuildEngine(configured); var selected = engine.SelectCategories(MigrationCategoryKind.Optional); var left = MigrationCategoryRegistry.All @@ -63,18 +64,27 @@ .. selected.Select(Describe), } [Test] - public void A_category_removed_from_configuration_leaves_its_checkpoint_row_untouched() + public async Task A_category_removed_from_configuration_leaves_its_checkpoint_row_untouched() { - var engine = BuildEngine([], out var checkpointStore); - var previousRun = new MigrationCheckpoint("EventLog", MigrationCategoryState.CompleteWithErrors, "cursor-99", 40, 2, 42, null, DateTime.UtcNow, DateTime.UtcNow, DateTime.UtcNow, null); - checkpointStore.Upsert(previousRun).GetAwaiter().GetResult(); + // Dropping a category from the configuration must not restart, reset or delete what it already + // copied: those rows are in the target, and a row wound back to the start copies every one again. + var stillConfigured = MigrationCategoryIds.EventLog; + var checkpointStore = new InMemoryMigrationCheckpointStore(); + var source = new InMemoryMigrationSource(); + source.Seed(stillConfigured, new MigrationRow("check-1", new object(), new Dictionary())); + var target = new InMemoryMigrationTarget(checkpointStore); + var options = new MigrationEngineOptions(TimeSpan.Zero, 5, 100, [stillConfigured]); + var engine = new MigrationEngine(source, target, checkpointStore, new FakeTimeProvider(), options, NullLogger.Instance); + var previousRun = new MigrationCheckpoint(MigrationCategoryIds.ArchivedAndResolvedFailedMessages, MigrationCategoryState.CompleteWithErrors, "cursor-99", 40, 2, 42, null, DateTime.UtcNow, DateTime.UtcNow, DateTime.UtcNow, null); + await checkpointStore.Upsert(previousRun); - var selected = engine.SelectCategories(MigrationCategoryKind.Optional); + var results = await engine.RunCategories(engine.SelectCategories(MigrationCategoryKind.Optional)); + var deselected = await checkpointStore.Read(MigrationCategoryIds.ArchivedAndResolvedFailedMessages); using (Assert.EnterMultipleScope()) { - Assert.That(selected.Select(c => c.Id), Does.Not.Contain("EventLog")); - Assert.That(checkpointStore.Read("EventLog").GetAwaiter().GetResult(), Is.EqualTo(previousRun with { Version = 1 })); + Assert.That(results.Select(c => c.CategoryId), Is.EqualTo(new[] { stillConfigured }), "the run touched only the configured category"); + Assert.That(deselected, Is.EqualTo(previousRun with { Version = 1 }), "the row is still at the version the seeding save left it"); } } } diff --git a/src/ServiceControl.UnitTests/Migration/MigrationEngineCopyTests.cs b/src/ServiceControl.UnitTests/Migration/MigrationEngineCopyTests.cs index aafb9d4935..6135b96aa5 100644 --- a/src/ServiceControl.UnitTests/Migration/MigrationEngineCopyTests.cs +++ b/src/ServiceControl.UnitTests/Migration/MigrationEngineCopyTests.cs @@ -40,6 +40,105 @@ public async Task Copies_every_row_and_finishes_Complete_when_nothing_was_skippe } } + [Test] + public async Task A_first_run_leaves_the_source_total_null_and_every_batch_moves_LastProgressAt() + { + var category = MigrationCategoryRegistry.Find("KnownEndpoints")!; + var source = new InMemoryMigrationSource(); + source.Seed(category.Id, Row("a"), Row("b"), Row("c"), Row("d")); + var checkpointStore = new InMemoryMigrationCheckpointStore(); + var target = new InMemoryMigrationTarget(checkpointStore) { DefaultBatchSize = 2 }; + var startedAt = new DateTimeOffset(2026, 3, 4, 5, 6, 7, TimeSpan.Zero); + var betweenBatches = TimeSpan.FromMinutes(1); + var clock = new FakeTimeProvider(startedAt); + var stamps = new List(); + // The engine stamps the checkpoint before the target sees it, so moving the clock in here is what + // makes a stamp that never moves show up as two identical entries. + target.BeforeWrite = extending => + { + stamps.Add(extending.LastProgressAt); + clock.Advance(betweenBatches); + }; + var options = new MigrationEngineOptions(TimeSpan.Zero, 5, 100, []); + var engine = new MigrationEngine(source, target, checkpointStore, clock, options, NullLogger.Instance); + + var checkpoint = await engine.RunCategoryAsync(category); + + using (Assert.EnterMultipleScope()) + { + Assert.That(checkpoint.SourceTotal, Is.Null, "counting the source is a full pass over it, which on a big category could trip the stall watchdog before a row is copied"); + Assert.That(stamps, Is.EqualTo(new DateTime?[] { startedAt.UtcDateTime, (startedAt + betweenBatches).UtcDateTime }), "the stall watchdog stops a copy whose LastProgressAt stops moving, so every batch has to move it"); + Assert.That(checkpoint.LastProgressAt, Is.EqualTo((startedAt + betweenBatches).UtcDateTime), "the saved row carries the last batch's stamp"); + } + } + + [Test] + public async Task A_category_that_read_fewer_rows_than_the_source_holds_still_settles_complete() + { + // The row says the source held 10 and the source yields 2. The engine judges only its own accounting, + // so a read that ends early is not something it can see, and only --migration-verify shows it. + var category = MigrationCategoryRegistry.Find("KnownEndpoints")!; + var source = new InMemoryMigrationSource(); + source.Seed(category.Id, Row("d"), Row("e")); + var checkpointStore = new InMemoryMigrationCheckpointStore(); + await checkpointStore.Upsert(new MigrationCheckpoint(category.Id, MigrationCategoryState.InProgress, null, 0, 0, 10, null, DateTime.UtcNow, DateTime.UtcNow, null, null)); + var target = new InMemoryMigrationTarget(checkpointStore); + var options = new MigrationEngineOptions(TimeSpan.Zero, 5, 100, []); + var engine = new MigrationEngine(source, target, checkpointStore, new FakeTimeProvider(), options, NullLogger.Instance); + + var checkpoint = await engine.RunCategoryAsync(category); + + using (Assert.EnterMultipleScope()) + { + Assert.That(checkpoint.State, Is.EqualTo(MigrationCategoryState.Complete)); + Assert.That(checkpoint.CopiedCount, Is.EqualTo(2)); + Assert.That(checkpoint.SettledAt, Is.Not.Null); + } + } + + [Test] + public async Task An_optional_category_whose_outcomes_do_not_add_up_to_the_rows_read_settles_halted() + { + var category = MigrationCategoryRegistry.Find(MigrationCategoryIds.EventLog)!; + var checkpoint = await RunWithOneRowUnaccountedFor(category); + + using (Assert.EnterMultipleScope()) + { + Assert.That(category.Kind, Is.EqualTo(MigrationCategoryKind.Optional)); + Assert.That(checkpoint.State, Is.EqualTo(MigrationCategoryState.Halted)); + Assert.That(checkpoint.LastError, Does.Contain("read 4 rows").And.Contain("accounted for 3")); + Assert.That(checkpoint.LastError, Does.Not.Contain("--migration-retry").And.Not.Contain("--migration-abandon")); + } + } + + [Test] + public async Task A_required_category_whose_outcomes_do_not_add_up_to_the_rows_read_settles_halted() + { + var category = MigrationCategoryRegistry.Find(MigrationCategoryIds.KnownEndpoints)!; + var checkpoint = await RunWithOneRowUnaccountedFor(category); + + using (Assert.EnterMultipleScope()) + { + Assert.That(category.Kind, Is.EqualTo(MigrationCategoryKind.Required)); + Assert.That(checkpoint.State, Is.EqualTo(MigrationCategoryState.Halted)); + Assert.That(checkpoint.LastError, Does.Contain("read 4 rows").And.Contain("accounted for 3")); + Assert.That(checkpoint.LastError, Does.Not.Contain("--migration-retry").And.Not.Contain("--migration-abandon")); + } + } + + // The target commits a checkpoint and a result that agree with each other, so only the rows read can show the one it lost. + static Task RunWithOneRowUnaccountedFor(MigrationCategory category) + { + var source = new InMemoryMigrationSource(); + source.Seed(category.Id, Row("a"), Row("b"), Row("c"), Row("d")); + var checkpointStore = new InMemoryMigrationCheckpointStore(); + var target = new InMemoryMigrationTarget(checkpointStore) { DefaultBatchSize = 4, UnderReportBy = 1 }; + var options = new MigrationEngineOptions(TimeSpan.Zero, 5, 100, []); + var engine = new MigrationEngine(source, target, checkpointStore, new FakeTimeProvider(), options, NullLogger.Instance); + + return engine.RunCategoryAsync(category); + } + [Test] public async Task A_category_with_no_rows_finishes_Complete_without_a_cursor() { @@ -84,8 +183,7 @@ public async Task The_moment_a_category_settles_comes_from_the_injected_clock() [Test] public async Task A_category_already_CompleteWithErrors_is_left_alone_on_a_second_run() { - // Finished with a few skips is finished. Re-reading it would copy the whole category again and - // count its skips a second time. + // Failed waits for the operator, so a start returns the row as it is rather than reading the category again. var category = MigrationCategoryRegistry.Find("KnownEndpoints")!; var source = new InMemoryMigrationSource(); source.Seed(category.Id, Row("a"), Row("b")); @@ -106,6 +204,29 @@ public async Task A_category_already_CompleteWithErrors_is_left_alone_on_a_secon } } + [Test] + public async Task A_category_already_Halted_is_left_alone_on_a_second_run() + { + // Failed waits for the operator, so a start returns the row as it is rather than reading the category again. + var category = MigrationCategoryRegistry.Find("KnownEndpoints")!; + var source = new InMemoryMigrationSource(); + source.Seed(category.Id, Row("a"), Row("b")); + var checkpointStore = new InMemoryMigrationCheckpointStore(); + var halted = new MigrationCheckpoint(category.Id, MigrationCategoryState.Halted, "a", 1, 0, null, null, DateTime.UtcNow, DateTime.UtcNow, DateTime.UtcNow, "Halted: the target was unreachable"); + await checkpointStore.Upsert(halted); + var target = new InMemoryMigrationTarget(checkpointStore); + var options = new MigrationEngineOptions(TimeSpan.Zero, 5, 100, []); + var engine = new MigrationEngine(source, target, checkpointStore, new FakeTimeProvider(), options, NullLogger.Instance); + + var checkpoint = await engine.RunCategoryAsync(category); + + using (Assert.EnterMultipleScope()) + { + Assert.That(checkpoint, Is.EqualTo(halted with { Version = 1 }), "the row is read back untouched, at the version the seeding save left it"); + Assert.That(target.WrittenRows(category.Id), Is.Empty); + } + } + [Test] public async Task A_category_already_Complete_is_left_alone_on_a_second_run() { diff --git a/src/ServiceControl.UnitTests/Migration/MigrationEngineFailurePathTests.cs b/src/ServiceControl.UnitTests/Migration/MigrationEngineFailurePathTests.cs index fb26628421..9d02e7bb4d 100644 --- a/src/ServiceControl.UnitTests/Migration/MigrationEngineFailurePathTests.cs +++ b/src/ServiceControl.UnitTests/Migration/MigrationEngineFailurePathTests.cs @@ -35,6 +35,7 @@ public async Task A_failing_write_halts_the_category_and_records_the_error_inste using (Assert.EnterMultipleScope()) { + Assert.That(category.Kind, Is.EqualTo(MigrationCategoryKind.Required)); Assert.That(checkpoint.State, Is.EqualTo(MigrationCategoryState.Halted)); Assert.That(checkpoint.LastError, Does.Contain("Simulated failure")); // The first batch committed, so the cursor is real and a later run resumes from it. @@ -78,40 +79,43 @@ public void A_cancelled_run_is_a_shutdown_rather_than_a_failure_and_is_not_swall } [Test] - public async Task A_halted_category_is_re_attempted_on_the_next_run_and_resumes_from_its_cursor() + public async Task A_halted_category_is_left_alone_until_the_operator_puts_it_back_in_progress() { - // A halt that no restart can clear would leave abandoning the category as the only way out of - // a fault the customer has already repaired. var category = MigrationCategoryRegistry.Find("KnownEndpoints")!; var source = new InMemoryMigrationSource(); source.Seed(category.Id, Row("a"), Row("b"), Row("c"), Row("d")); var checkpointStore = new InMemoryMigrationCheckpointStore(); var target = new InMemoryMigrationTarget(checkpointStore) { DefaultBatchSize = 2, FailOnCallNumber = 2 }; - var firstRun = BuildEngine(source, checkpointStore, target); - var halted = await firstRun.RunCategoryAsync(category); + var halted = await BuildEngine(source, checkpointStore, target).RunCategoryAsync(category); Assert.That(halted.State, Is.EqualTo(MigrationCategoryState.Halted)); - // Stopped on its first write, so the saved row is the restarted one rather than the completed one. + // The cause is fixed, but nothing re-runs a Failed category until the operator says so. target.FailOnCallNumber = null; - using var stopping = new CancellationTokenSource(); - target.StopOnCall = (3, stopping); - Assert.ThrowsAsync(() => BuildEngine(source, checkpointStore, target).RunCategoryAsync(category, stopping.Token)); - var restarted = await checkpointStore.Read(category.Id); + var untouched = await BuildEngine(source, checkpointStore, target).RunCategoryAsync(category); - target.StopOnCall = null; - var lastRun = BuildEngine(source, checkpointStore, target); - var finished = await lastRun.RunCategoryAsync(category); + // The reset --migration-retry writes: a re-read from the start that keeps only when the category first started. + await checkpointStore.Upsert(untouched with + { + State = MigrationCategoryState.NotStarted, + Cursor = null, + SkipReasons = null, + SettledAt = null, + LastError = null, + CopiedCount = 0, + SkippedCount = 0, + AlreadyPresentCount = 0 + }); + var finished = await BuildEngine(source, checkpointStore, target).RunCategoryAsync(category); using (Assert.EnterMultipleScope()) { - Assert.That(restarted!.State, Is.EqualTo(MigrationCategoryState.InProgress)); - Assert.That(restarted.SettledAt, Is.Null, "a copy running again does not keep the time its halt settled at"); + Assert.That(untouched, Is.EqualTo(halted), "a start re-ran a Failed category the operator had not put back"); Assert.That(finished.State, Is.EqualTo(MigrationCategoryState.Complete)); - Assert.That(finished.LastError, Is.Null, "a cleared halt does not leave a stale error on the row"); - Assert.That(finished.CopiedCount, Is.EqualTo(4)); - var writtenIds = target.WrittenRows(category.Id).Select(r => r.SourceId).ToArray(); - Assert.That(writtenIds, Is.EquivalentTo(new[] { "a", "b", "c", "d" })); - Assert.That(writtenIds.Distinct().Count(), Is.EqualTo(writtenIds.Length), "no duplicates across the halt"); + Assert.That(finished.LastError, Is.Null, "a retried category does not keep the error it failed with"); + Assert.That(finished.SourceTotal, Is.Null); + Assert.That(target.WrittenRows(category.Id).Select(r => r.SourceId), Is.EquivalentTo(new[] { "a", "b", "c", "d" })); + // Rows a and b were committed before the failure, so the re-read finds them already present. + Assert.That((finished.CopiedCount, finished.AlreadyPresentCount), Is.EqualTo((2L, 2L))); } } @@ -130,6 +134,8 @@ public async Task Body_skips_in_a_batch_whose_write_fails_are_counted_once_acros var firstRun = new MigrationEngine(source, target, checkpointStore, new FakeTimeProvider(), options, NullLogger.Instance); var halted = await firstRun.RunCategoryAsync(category); + // The row a crash just before the failing batch commits leaves, which a start resumes from its cursor. + await checkpointStore.Upsert(halted with { State = MigrationCategoryState.InProgress, LastError = null, SettledAt = null }); target.FailOnCallNumber = null; var secondRun = new MigrationEngine(source, target, checkpointStore, new FakeTimeProvider(), options, NullLogger.Instance); var finished = await secondRun.RunCategoryAsync(category); @@ -191,7 +197,7 @@ public async Task A_target_with_no_batch_size_for_a_category_halts_it_and_the_ne } [Test] - public void A_halt_whose_save_fails_still_logs_the_exception_that_caused_it() + public void A_write_failure_halt_logs_its_exception_before_it_settles() { var category = MigrationCategoryRegistry.Find("KnownEndpoints")!; var source = new InMemoryMigrationSource(); @@ -207,7 +213,7 @@ public void A_halt_whose_save_fails_still_logs_the_exception_that_caused_it() } [Test] - public void A_threshold_halt_whose_save_fails_still_logs_why_it_halted() + public void A_threshold_halt_logs_its_reason_before_it_settles() { var category = MigrationCategoryRegistry.Find("KnownEndpoints")!; var source = new InMemoryMigrationSource(); @@ -225,6 +231,78 @@ public void A_threshold_halt_whose_save_fails_still_logs_why_it_halted() Assert.That(logger.Entries.Where(e => e.Level == LogLevel.Error).Select(e => e.Message), Has.Some.Contains("Halted: 1 of 1 rows skipped")); } + [Test] + public void An_optional_threshold_halt_whose_save_fails_throws_rather_than_staying_copying() + { + var category = MigrationCategoryRegistry.Find(MigrationCategoryIds.EventLog)!; + var source = new InMemoryMigrationSource(); + source.Seed(category.Id, Row("a")); + var checkpointStore = new HaltSaveFailsCheckpointStore(); + var target = new InMemoryMigrationTarget(checkpointStore); + target.RejectKey("a", MigrationSkipReason.BodyUnreadable); + // A floor of zero lets the one rejected row halt the category. + var options = new MigrationEngineOptions(TimeSpan.Zero, HaltThresholdPercent: 5, HaltThresholdMinimum: 0, []); + var engine = new MigrationEngine(source, target, checkpointStore, new FakeTimeProvider(), options, NullLogger.Instance); + + Assert.ThrowsAsync(() => engine.RunCategoryAsync(category)); + + var stored = checkpointStore.Read(category.Id).GetAwaiter().GetResult(); + using (Assert.EnterMultipleScope()) + { + Assert.That(stored!.State, Is.EqualTo(MigrationCategoryState.InProgress)); + Assert.That(stored.LastError, Is.Null, "the failed save of a halt was recorded as the category's own error, which leaves a Failed category reading as still copying"); + } + } + + [Test] + public async Task An_optional_category_that_throws_stays_copying_with_the_error_and_its_cursor() + { + var category = MigrationCategoryRegistry.Find(MigrationCategoryIds.EventLog)!; + var source = new InMemoryMigrationSource(); + source.Seed(category.Id, Row("a"), Row("b"), Row("c"), Row("d")); + var checkpointStore = new InMemoryMigrationCheckpointStore(); + var target = new InMemoryMigrationTarget(checkpointStore) { DefaultBatchSize = 2, FailOnCallNumber = 2 }; + + var checkpoint = await BuildEngine(source, checkpointStore, target).RunCategoryAsync(category); + + var stored = await checkpointStore.Read(category.Id); + using (Assert.EnterMultipleScope()) + { + Assert.That(category.Kind, Is.EqualTo(MigrationCategoryKind.Optional)); + foreach (var row in new[] { checkpoint, stored! }) + { + Assert.That(row.State, Is.EqualTo(MigrationCategoryState.InProgress)); + Assert.That(row.Cursor, Is.EqualTo("b")); + Assert.That(row.CopiedCount, Is.EqualTo(2)); + Assert.That(row.SettledAt, Is.Null); + Assert.That(row.LastError, Does.Contain("Simulated failure")); + Assert.That(row.LastError, Does.Not.Contain("--migration-abandon").And.Not.Contain("--migration-retry")); + } + } + } + + [Test] + public async Task An_optional_category_an_exception_left_copying_resumes_on_the_next_start_and_clears_its_error() + { + var category = MigrationCategoryRegistry.Find(MigrationCategoryIds.EventLog)!; + var source = new InMemoryMigrationSource(); + source.Seed(category.Id, Row("a"), Row("b"), Row("c"), Row("d")); + var checkpointStore = new InMemoryMigrationCheckpointStore(); + var target = new InMemoryMigrationTarget(checkpointStore) { DefaultBatchSize = 2, FailOnCallNumber = 2 }; + await BuildEngine(source, checkpointStore, target).RunCategoryAsync(category); + + target.FailOnCallNumber = null; + var finished = await BuildEngine(source, checkpointStore, target).RunCategoryAsync(category); + + using (Assert.EnterMultipleScope()) + { + Assert.That(finished.State, Is.EqualTo(MigrationCategoryState.Complete)); + Assert.That(finished.LastError, Is.Null, "the error stays on the row only until a start tries again"); + Assert.That(finished.CopiedCount, Is.EqualTo(4)); + Assert.That(target.RowsHandedToWrite(category.Id).Select(r => r.SourceId), Is.EqualTo(new[] { "a", "b", "c", "d" }), "the resume did not carry on from the cursor"); + } + } + [Test] public async Task A_checkpoint_conflict_leaves_the_other_writer_alone_instead_of_halting_over_it() { diff --git a/src/ServiceControl.UnitTests/Migration/MigrationEngineHaltTests.cs b/src/ServiceControl.UnitTests/Migration/MigrationEngineHaltTests.cs index 47a77f85d7..0ece00a601 100644 --- a/src/ServiceControl.UnitTests/Migration/MigrationEngineHaltTests.cs +++ b/src/ServiceControl.UnitTests/Migration/MigrationEngineHaltTests.cs @@ -58,7 +58,7 @@ public async Task Rows_the_target_would_have_deleted_anyway_never_count_toward_t var target = new InMemoryMigrationTarget(checkpointStore) { DefaultBatchSize = 100 }; foreach (var i in Enumerable.Range(1, 1_000).Where(i => i % 5 == 0)) { - target.RejectKey($"row-{i}", MigrationSkipReason.PastRetention, benign: true); + target.RejectKey($"row-{i}", MigrationSkipReason.PastRetention); } var options = new MigrationEngineOptions(TimeSpan.Zero, HaltThresholdPercent: 5, HaltThresholdMinimum: 100, []); var engine = new MigrationEngine(source, target, checkpointStore, new FakeTimeProvider(), options, NullLogger.Instance); @@ -67,7 +67,7 @@ public async Task Rows_the_target_would_have_deleted_anyway_never_count_toward_t using (Assert.EnterMultipleScope()) { - Assert.That(checkpoint.State, Is.EqualTo(MigrationCategoryState.CompleteWithErrors)); + Assert.That(checkpoint.State, Is.EqualTo(MigrationCategoryState.Complete)); Assert.That((checkpoint.CopiedCount, checkpoint.SkippedCount), Is.EqualTo((800L, 200L))); Assert.That(target.WrittenRows(category.Id), Has.Count.EqualTo(800), "the last row was reached, so nothing halted partway"); } @@ -78,7 +78,7 @@ public async Task Rows_the_target_would_have_deleted_anyway_never_count_toward_t public async Task A_batch_mixing_benign_and_fault_skips_is_judged_on_the_faults_alone(int everyNthIsAFault, MigrationCategoryState expected) { // A real archive copy loses rows both ways at once: retention takes some, unreadable bodies take - // others. This is the only shape where the subtraction has to do arithmetic rather than pick a side. + // others. This is the only shape where the sum has to leave some reasons out. var category = MigrationCategoryRegistry.Find("ArchivedAndResolvedFailedMessages")!; var source = new InMemoryMigrationSource(); source.Seed(category.Id, [.. Enumerable.Range(1, 5_000).Select(i => Row($"row-{i}"))]); @@ -87,7 +87,7 @@ public async Task A_batch_mixing_benign_and_fault_skips_is_judged_on_the_faults_ // A fifth of the category is past retention either way, which on its own is four times the threshold. foreach (var i in Enumerable.Range(1, 5_000).Where(i => i % 5 == 0)) { - target.RejectKey($"row-{i}", MigrationSkipReason.PastRetention, benign: true); + target.RejectKey($"row-{i}", MigrationSkipReason.PastRetention); } // The offset keeps the faults clear of the benign rows: 4% of the category in one case, 10% in the other. foreach (var i in Enumerable.Range(1, 5_000).Where(i => i % everyNthIsAFault == 3)) @@ -160,30 +160,167 @@ public async Task Rows_already_present_in_the_target_never_count_toward_the_halt } [Test] - public async Task A_restart_after_a_threshold_halt_counts_only_its_own_skips_and_keeps_the_earlier_ones() + public async Task A_resumed_run_counts_only_its_own_skips_and_keeps_the_earlier_ones() { var category = MigrationCategoryRegistry.Find("KnownEndpoints")!; var source = new InMemoryMigrationSource(); source.Seed(category.Id, [.. Enumerable.Range(1, 1_000).Select(i => Row($"row-{i}"))]); var checkpointStore = new InMemoryMigrationCheckpointStore(); - var failingTarget = new InMemoryMigrationTarget(checkpointStore) { DefaultBatchSize = 100 }; + // What a shutdown after the sixth batch leaves: 120 fault skips in 600 rows, past both the floor and 5% if counted again. + await checkpointStore.Upsert(new MigrationCheckpoint(category.Id, MigrationCategoryState.InProgress, "row-600", 480, 120, null, + new Dictionary { [MigrationSkipReason.BodyUnreadable] = 120 }, DateTime.UtcNow, DateTime.UtcNow, null, null)); + var options = new MigrationEngineOptions(TimeSpan.Zero, HaltThresholdPercent: 5, HaltThresholdMinimum: 100, []); + var target = new InMemoryMigrationTarget(checkpointStore) { DefaultBatchSize = 100 }; + + var finished = await new MigrationEngine(source, target, checkpointStore, new FakeTimeProvider(), options, NullLogger.Instance).RunCategoryAsync(category); + + using (Assert.EnterMultipleScope()) + { + Assert.That(finished.State, Is.EqualTo(MigrationCategoryState.CompleteWithErrors), "the skips still on the row must not halt a run that skips nothing"); + Assert.That((finished.CopiedCount, finished.SkippedCount), Is.EqualTo((880L, 120L)), "copied, skipped at the end"); + } + } + + [Test] + public async Task A_category_smaller_than_the_floor_that_loses_every_row_ends_Failed_rather_than_Done() + { + // Ninety rows is under the hundred-row floor, so the threshold the engine checks after every batch + // can never fire, however many rows are lost. + var category = MigrationCategoryRegistry.Find(MigrationCategoryIds.KnownEndpoints)!; + var source = new InMemoryMigrationSource(); + source.Seed(category.Id, [.. Enumerable.Range(1, 90).Select(i => Row($"row-{i}"))]); + var checkpointStore = new InMemoryMigrationCheckpointStore(); + var target = new InMemoryMigrationTarget(checkpointStore) { DefaultBatchSize = 30 }; + foreach (var i in Enumerable.Range(1, 90)) + { + target.RejectKey($"row-{i}", MigrationSkipReason.RequiredValueMissing); + } + var options = new MigrationEngineOptions(TimeSpan.Zero, HaltThresholdPercent: 5, HaltThresholdMinimum: 100, []); + + var checkpoint = await new MigrationEngine(source, target, checkpointStore, new FakeTimeProvider(), options, NullLogger.Instance).RunCategoryAsync(category); + + using (Assert.EnterMultipleScope()) + { + Assert.That(checkpoint.State, Is.EqualTo(MigrationCategoryState.CompleteWithErrors)); + Assert.That(checkpoint.State.IsFinished(), Is.False, "a required category that copied nothing must not let the host open"); + Assert.That(checkpoint.CopiedCount, Is.Zero); + } + } + + [Test] + public async Task A_halted_category_is_returned_untouched_and_its_source_is_not_read() + { + // Failed waits for the operator, so a start that read the category again would copy rows nobody asked for. + var category = MigrationCategoryRegistry.Find(MigrationCategoryIds.KnownEndpoints)!; + var source = new InMemoryMigrationSource(); + source.Seed(category.Id, [.. Enumerable.Range(1, 90).Select(i => Row($"row-{i}"))]); + var checkpointStore = new InMemoryMigrationCheckpointStore(); + var halted = new MigrationCheckpoint(category.Id, MigrationCategoryState.Halted, null, 0, 0, null, null, DateTime.UtcNow, DateTime.UtcNow, DateTime.UtcNow, "Halted: the target was unreachable"); + await checkpointStore.Upsert(halted); + var target = new InMemoryMigrationTarget(checkpointStore) { DefaultBatchSize = 30 }; + var options = new MigrationEngineOptions(TimeSpan.Zero, HaltThresholdPercent: 5, HaltThresholdMinimum: 100, []); + var engine = new MigrationEngine(source, target, checkpointStore, new FakeTimeProvider(), options, NullLogger.Instance); + + var checkpoint = await engine.RunCategoryAsync(category); + + using (Assert.EnterMultipleScope()) + { + Assert.That(checkpoint, Is.EqualTo(halted with { Version = 1 }), "the row is read back untouched, at the version the seeding save left it"); + Assert.That(target.RowsHandedToWrite(category.Id), Is.Empty); + } + } + + [Test] + public async Task A_small_category_losing_rows_the_product_would_drop_anyway_still_completes() + { + // Harmless skips are rows the target would have deleted anyway, so no number of them may stop a + // category or leave it Failed, however small the category is. + var category = MigrationCategoryRegistry.Find(MigrationCategoryIds.KnownEndpoints)!; + var source = new InMemoryMigrationSource(); + source.Seed(category.Id, [.. Enumerable.Range(1, 90).Select(i => Row($"row-{i}"))]); + var checkpointStore = new InMemoryMigrationCheckpointStore(); + var target = new InMemoryMigrationTarget(checkpointStore) { DefaultBatchSize = 30 }; + foreach (var i in Enumerable.Range(1, 90)) + { + target.RejectKey($"row-{i}", MigrationSkipReason.PastRetention); + } + var options = new MigrationEngineOptions(TimeSpan.Zero, HaltThresholdPercent: 5, HaltThresholdMinimum: 100, []); + + var checkpoint = await new MigrationEngine(source, target, checkpointStore, new FakeTimeProvider(), options, NullLogger.Instance).RunCategoryAsync(category); + + Assert.That(checkpoint.State, Is.EqualTo(MigrationCategoryState.Complete)); + } + + [Test] + public async Task A_run_that_ends_with_one_fault_skip_settles_CompleteWithErrors_however_small_the_share() + { + var category = MigrationCategoryRegistry.Find(MigrationCategoryIds.KnownEndpoints)!; + var source = new InMemoryMigrationSource(); + source.Seed(category.Id, [.. Enumerable.Range(1, 1_000).Select(i => Row($"row-{i}"))]); + var checkpointStore = new InMemoryMigrationCheckpointStore(); + var target = new InMemoryMigrationTarget(checkpointStore) { DefaultBatchSize = 100 }; + target.RejectKey("row-500", MigrationSkipReason.RequiredValueMissing); + var options = new MigrationEngineOptions(TimeSpan.Zero, HaltThresholdPercent: 5, HaltThresholdMinimum: 100, []); + + var checkpoint = await new MigrationEngine(source, target, checkpointStore, new FakeTimeProvider(), options, NullLogger.Instance).RunCategoryAsync(category); + + using (Assert.EnterMultipleScope()) + { + Assert.That(checkpoint.State, Is.EqualTo(MigrationCategoryState.CompleteWithErrors), "one row in a thousand is far under the threshold, and it is still a row the product wanted"); + Assert.That((checkpoint.CopiedCount, checkpoint.SkippedCount), Is.EqualTo((999L, 1L))); + } + } + + [Test] + public async Task A_run_whose_only_skips_are_harmless_settles_Complete_and_keeps_its_reasons() + { + var category = MigrationCategoryRegistry.Find(MigrationCategoryIds.KnownEndpoints)!; + var source = new InMemoryMigrationSource(); + source.Seed(category.Id, Row("a"), Row("b"), Row("c"), Row("d")); + var checkpointStore = new InMemoryMigrationCheckpointStore(); + var target = new InMemoryMigrationTarget(checkpointStore) { DefaultBatchSize = 2 }; + target.RejectKey("a", MigrationSkipReason.PastRetention); + target.RejectKey("b", MigrationSkipReason.PastRetention); + target.RejectKey("c", MigrationSkipReason.BlankGroupComment); + var options = new MigrationEngineOptions(TimeSpan.Zero, HaltThresholdPercent: 5, HaltThresholdMinimum: 100, []); + + var checkpoint = await new MigrationEngine(source, target, checkpointStore, new FakeTimeProvider(), options, NullLogger.Instance).RunCategoryAsync(category); + + using (Assert.EnterMultipleScope()) + { + Assert.That(checkpoint.State, Is.EqualTo(MigrationCategoryState.Complete)); + Assert.That((checkpoint.CopiedCount, checkpoint.SkippedCount), Is.EqualTo((1L, 3L)), "a harmless skip is still a skip, and status and verify count it"); + Assert.That(checkpoint.SkipReasons, Is.EquivalentTo(new Dictionary + { + [MigrationSkipReason.PastRetention] = 2, + [MigrationSkipReason.BlankGroupComment] = 1 + })); + } + } + + [Test] + public async Task The_threshold_halt_names_its_counts_and_no_command() + { + var category = MigrationCategoryRegistry.Find(MigrationCategoryIds.ArchivedAndResolvedFailedMessages)!; + var source = new InMemoryMigrationSource(); + source.Seed(category.Id, [.. Enumerable.Range(1, 1_000).Select(i => Row($"row-{i}"))]); + var checkpointStore = new InMemoryMigrationCheckpointStore(); + var target = new InMemoryMigrationTarget(checkpointStore) { DefaultBatchSize = 100 }; foreach (var i in Enumerable.Range(1, 1_000).Where(i => i % 5 == 0)) { - failingTarget.RejectKey($"row-{i}", MigrationSkipReason.BodyUnreadable); + target.RejectKey($"row-{i}", MigrationSkipReason.BodyUnreadable); } var options = new MigrationEngineOptions(TimeSpan.Zero, HaltThresholdPercent: 5, HaltThresholdMinimum: 100, []); - var halted = await new MigrationEngine(source, failingTarget, checkpointStore, new FakeTimeProvider(), options, NullLogger.Instance).RunCategoryAsync(category); - // The cause is fixed: the remaining rows now write cleanly. - var fixedTarget = new InMemoryMigrationTarget(checkpointStore) { DefaultBatchSize = 100 }; - var finished = await new MigrationEngine(source, fixedTarget, checkpointStore, new FakeTimeProvider(), options, NullLogger.Instance).RunCategoryAsync(category); + var checkpoint = await new MigrationEngine(source, target, checkpointStore, new FakeTimeProvider(), options, NullLogger.Instance).RunCategoryAsync(category); using (Assert.EnterMultipleScope()) { + Assert.That(checkpoint.State, Is.EqualTo(MigrationCategoryState.Halted)); // 120 skips in 600 rows is the first point past both the floor and 5%. - Assert.That((halted.State, halted.CopiedCount, halted.SkippedCount), Is.EqualTo((MigrationCategoryState.Halted, 480L, 120L))); - Assert.That(finished.State, Is.EqualTo(MigrationCategoryState.CompleteWithErrors), "the skips still on the row must not halt a run that skips nothing"); - Assert.That((finished.CopiedCount, finished.SkippedCount), Is.EqualTo((880L, 120L)), "copied, skipped at the end"); + Assert.That(checkpoint.LastError, Does.Contain("120 of 600").And.Contain("5%").And.Contain("100 rows")); + // The engine cannot know which commands a category may take, so the readers of the row add them. + Assert.That(checkpoint.LastError, Does.Not.Contain("--migration-retry").And.Not.Contain("--migration-abandon").And.Not.Contain("restart to resume")); } } } diff --git a/src/ServiceControl.UnitTests/Migration/MigrationEngineOptionsTests.cs b/src/ServiceControl.UnitTests/Migration/MigrationEngineOptionsTests.cs index 0ede2f09b3..71e62dac7b 100644 --- a/src/ServiceControl.UnitTests/Migration/MigrationEngineOptionsTests.cs +++ b/src/ServiceControl.UnitTests/Migration/MigrationEngineOptionsTests.cs @@ -10,6 +10,8 @@ namespace ServiceControl.UnitTests.Migration; class MigrationEngineOptionsTests { static readonly SettingsRootNamespace Namespace = new("ServiceControl"); + static readonly TimeSpan EventRetention = TimeSpan.FromDays(3); + static readonly TimeSpan ErrorRetention = TimeSpan.FromDays(11); [TearDown] public void ClearEnvironmentVariables() @@ -17,20 +19,25 @@ public void ClearEnvironmentVariables() Environment.SetEnvironmentVariable("SERVICECONTROL_MIGRATION_THROTTLEPAUSEMILLISECONDS", null); Environment.SetEnvironmentVariable("SERVICECONTROL_MIGRATION_HALTTHRESHOLDPERCENT", null); Environment.SetEnvironmentVariable("SERVICECONTROL_MIGRATION_HALTTHRESHOLDMINIMUM", null); - Environment.SetEnvironmentVariable("SERVICECONTROL_MIGRATION_OPTIONALCATEGORIES", null); + Environment.SetEnvironmentVariable("SERVICECONTROL_MIGRATION_EVENTLOGWINDOW", null); + Environment.SetEnvironmentVariable("SERVICECONTROL_MIGRATION_ARCHIVEDANDRESOLVEDFAILEDMESSAGESWINDOW", null); } + static MigrationEngineOptions Read() => MigrationEngineOptions.FromSettings(Namespace, EventRetention, ErrorRetention); + [Test] public void Defaults_match_the_contract_when_nothing_is_configured() { - var options = MigrationEngineOptions.FromSettings(Namespace); + var options = Read(); using (Assert.EnterMultipleScope()) { Assert.That(options.ThrottlePause, Is.EqualTo(TimeSpan.FromMilliseconds(100))); Assert.That(options.HaltThresholdPercent, Is.EqualTo(5)); Assert.That(options.HaltThresholdMinimum, Is.EqualTo(100)); - Assert.That(options.SelectedOptionalCategoryIds, Is.Empty); + Assert.That(options.SelectedOptionalCategoryIds, Is.EquivalentTo(new[] { MigrationCategoryIds.EventLog, MigrationCategoryIds.ArchivedAndResolvedFailedMessages }), "both optional categories are copied unless turned off"); + Assert.That(options.EventLogWindow, Is.EqualTo(EventRetention), "the event log window defaults to the event retention period"); + Assert.That(options.ArchivedAndResolvedFailedMessagesWindow, Is.EqualTo(ErrorRetention), "the archived and resolved window defaults to the error retention period"); Assert.That(options.BodyRetryBackoff, Is.EqualTo(TimeSpan.FromMilliseconds(200))); } } @@ -41,38 +48,41 @@ public void Reads_configured_values_from_environment_variables() Environment.SetEnvironmentVariable("SERVICECONTROL_MIGRATION_THROTTLEPAUSEMILLISECONDS", "2500"); Environment.SetEnvironmentVariable("SERVICECONTROL_MIGRATION_HALTTHRESHOLDPERCENT", "10"); Environment.SetEnvironmentVariable("SERVICECONTROL_MIGRATION_HALTTHRESHOLDMINIMUM", "50"); - Environment.SetEnvironmentVariable("SERVICECONTROL_MIGRATION_OPTIONALCATEGORIES", "EventLog, CustomChecks"); + Environment.SetEnvironmentVariable("SERVICECONTROL_MIGRATION_EVENTLOGWINDOW", "2.00:00:00"); + Environment.SetEnvironmentVariable("SERVICECONTROL_MIGRATION_ARCHIVEDANDRESOLVEDFAILEDMESSAGESWINDOW", "5.12:00:00"); - var options = MigrationEngineOptions.FromSettings(Namespace); + var options = Read(); using (Assert.EnterMultipleScope()) { Assert.That(options.ThrottlePause, Is.EqualTo(TimeSpan.FromMilliseconds(2500))); Assert.That(options.HaltThresholdPercent, Is.EqualTo(10)); Assert.That(options.HaltThresholdMinimum, Is.EqualTo(50)); - Assert.That(options.SelectedOptionalCategoryIds, Is.EquivalentTo(new[] { "EventLog", "CustomChecks" })); + Assert.That(options.EventLogWindow, Is.EqualTo(TimeSpan.FromDays(2))); + Assert.That(options.ArchivedAndResolvedFailedMessagesWindow, Is.EqualTo(TimeSpan.FromDays(5.5))); + Assert.That(options.SelectedOptionalCategoryIds, Is.EquivalentTo(new[] { MigrationCategoryIds.EventLog, MigrationCategoryIds.ArchivedAndResolvedFailedMessages })); } } - [Test] - public void Refuses_an_unknown_optional_category_id() + [TestCase("SERVICECONTROL_MIGRATION_EVENTLOGWINDOW", MigrationCategoryIds.ArchivedAndResolvedFailedMessages)] + [TestCase("SERVICECONTROL_MIGRATION_ARCHIVEDANDRESOLVEDFAILEDMESSAGESWINDOW", MigrationCategoryIds.EventLog)] + public void A_zero_window_turns_its_category_off(string variable, string stillSelected) { - Environment.SetEnvironmentVariable("SERVICECONTROL_MIGRATION_OPTIONALCATEGORIES", "NoSuchCategory"); + Environment.SetEnvironmentVariable(variable, "0"); - var ex = Assert.Throws(() => MigrationEngineOptions.FromSettings(Namespace)); + var options = Read(); - Assert.That(ex.Message, Does.Contain("NoSuchCategory")); + Assert.That(options.SelectedOptionalCategoryIds, Is.EquivalentTo(new[] { stillSelected })); } - [Test] - public void Refuses_a_required_category_id_named_as_optional() + [TestCase("a week")] + [TestCase("-1.00:00:00")] + public void Refuses_a_window_that_is_not_a_time_span_of_zero_or_more(string value) { - // EndpointSettings is required, not optional: naming it here is a customer mistake, not - // a valid way to force it. Required categories are never a matter of configuration. - Environment.SetEnvironmentVariable("SERVICECONTROL_MIGRATION_OPTIONALCATEGORIES", "EndpointSettings"); + Environment.SetEnvironmentVariable("SERVICECONTROL_MIGRATION_EVENTLOGWINDOW", value); - var ex = Assert.Throws(() => MigrationEngineOptions.FromSettings(Namespace)); + var ex = Assert.Throws(() => Read()); - Assert.That(ex.Message, Does.Contain("EndpointSettings")); + Assert.That(ex.Message, Does.Contain(value).And.Contain(MigrationSettings.EventLogWindowKey), "the refusal has to name both the value and the setting holding it for the customer to fix it"); } } diff --git a/src/ServiceControl.UnitTests/Migration/MigrationEngineOrderingTests.cs b/src/ServiceControl.UnitTests/Migration/MigrationEngineOrderingTests.cs index a7f2c99e25..056805e938 100644 --- a/src/ServiceControl.UnitTests/Migration/MigrationEngineOrderingTests.cs +++ b/src/ServiceControl.UnitTests/Migration/MigrationEngineOrderingTests.cs @@ -67,13 +67,14 @@ public async Task EndpointSettings_does_not_start_before_KnownEndpoints_complete [TestCase(MigrationCategoryState.InProgress)] [TestCase(MigrationCategoryState.Halted)] - public async Task GroupComments_does_not_start_before_the_archive_completes(MigrationCategoryState archiveState) + [TestCase(MigrationCategoryState.CompleteWithErrors)] + public async Task GroupComments_does_not_start_before_the_unresolved_failed_messages_complete(MigrationCategoryState predecessorState) { var comments = MigrationCategoryRegistry.Find("GroupComments")!; var source = new InMemoryMigrationSource(); source.Seed(comments.Id, Row("GroupComment/g-1")); var checkpointStore = new InMemoryMigrationCheckpointStore(); - await checkpointStore.Upsert(new MigrationCheckpoint("ArchivedAndResolvedFailedMessages", archiveState, "m-500", 500, 0, null, null, DateTime.UtcNow, DateTime.UtcNow, null, null)); + await checkpointStore.Upsert(new MigrationCheckpoint("UnresolvedAndRetryIssuedFailedMessages", predecessorState, "m-500", 500, 0, null, null, DateTime.UtcNow, DateTime.UtcNow, null, null)); var target = new InMemoryMigrationTarget(checkpointStore); var engine = BuildEngine(source, checkpointStore, target); @@ -82,7 +83,7 @@ public async Task GroupComments_does_not_start_before_the_archive_completes(Migr using (Assert.EnterMultipleScope()) { Assert.That(checkpoint.State, Is.EqualTo(MigrationCategoryState.Blocked)); - Assert.That(checkpoint.LastError, Is.EqualTo($"Blocked: GroupComments must follow ArchivedAndResolvedFailedMessages, which is {archiveState}")); + Assert.That(checkpoint.LastError, Is.EqualTo($"Blocked: GroupComments must follow UnresolvedAndRetryIssuedFailedMessages, which is {predecessorState}")); Assert.That(target.WrittenRows(comments.Id), Is.Empty); } } @@ -90,15 +91,14 @@ public async Task GroupComments_does_not_start_before_the_archive_completes(Migr [Test] public async Task A_blocked_category_runs_once_the_category_it_follows_settles() { - // Every real migration starts group comments blocked, so a block nothing can clear would strand - // the last category and leave the migration unable to end. + // A block nothing can clear would strand group comments and keep the host closed for good. var comments = MigrationCategoryRegistry.Find("GroupComments")!; - var archive = MigrationCategoryRegistry.Find("ArchivedAndResolvedFailedMessages")!; + var unresolved = MigrationCategoryRegistry.Find("UnresolvedAndRetryIssuedFailedMessages")!; var source = new InMemoryMigrationSource(); source.Seed(comments.Id, Row("GroupComment/g-1")); var checkpointStore = new InMemoryMigrationCheckpointStore(); - var archiveRunning = new MigrationCheckpoint(archive.Id, MigrationCategoryState.InProgress, "m-500", 500, 0, null, null, DateTime.UtcNow, DateTime.UtcNow, null, null); - var saved = await checkpointStore.Upsert(archiveRunning); + var predecessorRunning = new MigrationCheckpoint(unresolved.Id, MigrationCategoryState.InProgress, "m-500", 500, 0, null, null, DateTime.UtcNow, DateTime.UtcNow, null, null); + var saved = await checkpointStore.Upsert(predecessorRunning); var target = new InMemoryMigrationTarget(checkpointStore); var blocked = await BuildEngine(source, checkpointStore, target).RunCategoryAsync(comments); @@ -133,7 +133,7 @@ public async Task A_blocked_category_has_not_settled_because_it_has_not_finished [TestCase(MigrationCategoryState.NotStarted)] [TestCase(MigrationCategoryState.Blocked)] - public async Task A_predecessor_that_has_a_row_but_has_not_run_is_named_by_the_state_on_that_row(MigrationCategoryState archiveState) + public async Task A_predecessor_that_has_a_row_but_has_not_run_is_named_by_the_state_on_that_row(MigrationCategoryState predecessorState) { // A missing row reads as "not started"; a row that exists says what it actually holds, which is // how an operator tells a category waiting its turn from one waiting on a chain. @@ -141,7 +141,7 @@ public async Task A_predecessor_that_has_a_row_but_has_not_run_is_named_by_the_s var source = new InMemoryMigrationSource(); source.Seed(comments.Id, Row("GroupComment/g-1")); var checkpointStore = new InMemoryMigrationCheckpointStore(); - await checkpointStore.Upsert(new MigrationCheckpoint("ArchivedAndResolvedFailedMessages", archiveState, null, 0, 0, null, null, null, null, null, null)); + await checkpointStore.Upsert(new MigrationCheckpoint("UnresolvedAndRetryIssuedFailedMessages", predecessorState, null, 0, 0, null, null, null, null, null, null)); var target = new InMemoryMigrationTarget(checkpointStore); var checkpoint = await BuildEngine(source, checkpointStore, target).RunCategoryAsync(comments); @@ -149,12 +149,11 @@ public async Task A_predecessor_that_has_a_row_but_has_not_run_is_named_by_the_s using (Assert.EnterMultipleScope()) { Assert.That(checkpoint.State, Is.EqualTo(MigrationCategoryState.Blocked)); - Assert.That(checkpoint.LastError, Is.EqualTo($"Blocked: GroupComments must follow ArchivedAndResolvedFailedMessages, which is {archiveState}")); + Assert.That(checkpoint.LastError, Is.EqualTo($"Blocked: GroupComments must follow UnresolvedAndRetryIssuedFailedMessages, which is {predecessorState}")); } } [TestCase(MigrationCategoryState.Complete)] - [TestCase(MigrationCategoryState.CompleteWithErrors)] [TestCase(MigrationCategoryState.Abandoned)] public async Task LicensingThroughput_proceeds_once_LicensingEndpoints_is_finished_or_abandoned(MigrationCategoryState endpointsState) { diff --git a/src/ServiceControl.UnitTests/Migration/MigrationEngineResumeTests.cs b/src/ServiceControl.UnitTests/Migration/MigrationEngineResumeTests.cs index d21760171f..3a2dca44d7 100644 --- a/src/ServiceControl.UnitTests/Migration/MigrationEngineResumeTests.cs +++ b/src/ServiceControl.UnitTests/Migration/MigrationEngineResumeTests.cs @@ -32,8 +32,19 @@ public async Task A_restart_keeps_the_moment_the_category_first_started() var halted = await new MigrationEngine(source, target, checkpointStore, new FakeTimeProvider(firstStart), options, NullLogger.Instance) .RunCategoryAsync(category); - // A day later, the cause is fixed and the host is started again. + // A day later, the cause is fixed and the operator puts the category back as --migration-retry does. target.FailOnCallNumber = null; + await checkpointStore.Upsert(halted with + { + State = MigrationCategoryState.NotStarted, + Cursor = null, + SkipReasons = null, + SettledAt = null, + LastError = null, + CopiedCount = 0, + SkippedCount = 0, + AlreadyPresentCount = 0 + }); var restart = firstStart.AddDays(1); var finished = await new MigrationEngine(source, target, checkpointStore, new FakeTimeProvider(restart), options, NullLogger.Instance) .RunCategoryAsync(category); @@ -42,7 +53,7 @@ public async Task A_restart_keeps_the_moment_the_category_first_started() { Assert.That(halted.StartedAt, Is.EqualTo(firstStart.UtcDateTime)); Assert.That(finished.State, Is.EqualTo(MigrationCategoryState.Complete)); - Assert.That(finished.StartedAt, Is.EqualTo(firstStart.UtcDateTime), "the restart carries on a copy that started a day ago"); + Assert.That(finished.StartedAt, Is.EqualTo(firstStart.UtcDateTime), "a retried category keeps the moment it first started"); Assert.That(finished.SettledAt, Is.EqualTo(restart.UtcDateTime)); } } @@ -82,9 +93,9 @@ public async Task Restarting_after_a_mid_category_stop_produces_no_duplicates_an { Assert.That(finalCheckpoint.State, Is.EqualTo(MigrationCategoryState.Complete)); Assert.That(finalCheckpoint.CopiedCount, Is.EqualTo(6)); - var writtenIds = target.WrittenRows(category.Id).Select(r => r.SourceId).ToArray(); - Assert.That(writtenIds, Is.EquivalentTo(allIds), "no gaps"); - Assert.That(writtenIds.Distinct().Count(), Is.EqualTo(writtenIds.Length), "no duplicates"); + Assert.That(target.WrittenRows(category.Id).Select(r => r.SourceId), Is.EquivalentTo(allIds), "no gaps"); + // The target de-duplicates, as the real ones do, so what it kept can never show a row sent twice. + Assert.That(target.RowsHandedToWrite(category.Id).Select(r => r.SourceId), Is.Unique, "no duplicates: the restart resumes past the rows the first run committed"); } } diff --git a/src/ServiceControl.UnitTests/Migration/MigrationEngineRunCategoriesTests.cs b/src/ServiceControl.UnitTests/Migration/MigrationEngineRunCategoriesTests.cs index e4e8393de0..d08b17f548 100644 --- a/src/ServiceControl.UnitTests/Migration/MigrationEngineRunCategoriesTests.cs +++ b/src/ServiceControl.UnitTests/Migration/MigrationEngineRunCategoriesTests.cs @@ -40,6 +40,80 @@ public async Task Runs_every_category_it_is_given_in_the_order_it_is_given_them( } } + [Test] + public async Task Every_category_it_is_given_has_a_row_before_the_first_one_copies_anything() + { + // A copy stopped between two categories must not leave a table that reads as finished, because the + // gates that keep a host off an unfinished copy look only at the rows that exist. + var source = new InMemoryMigrationSource(); + var checkpointStore = new InMemoryMigrationCheckpointStore(); + IReadOnlyList? rowsAtFirstWrite = null; + var target = new InMemoryMigrationTarget(checkpointStore) { BeforeWrite = _ => rowsAtFirstWrite ??= checkpointStore.ReadAll().GetAwaiter().GetResult() }; + var engine = new MigrationEngine(source, target, checkpointStore, new FakeTimeProvider(), + new MigrationEngineOptions(TimeSpan.Zero, 5, 100, []), NullLogger.Instance); + // Three, so recording only the next category would still leave the last one missing. + MigrationCategory[] given = [MigrationCategoryRegistry.Find("KnownEndpoints")!, MigrationCategoryRegistry.Find("EndpointSettings")!, MigrationCategoryRegistry.Find("MessageRedirects")!]; + foreach (var category in given) + { + source.Seed(category.Id, Row($"{category.Id}-1")); + } + + await engine.RunCategories(given); + + Assert.That(rowsAtFirstWrite, Is.Not.Null, "the copy never wrote a batch"); + using (Assert.EnterMultipleScope()) + { + Assert.That(rowsAtFirstWrite!.Select(c => c.CategoryId), Is.EquivalentTo(new[] { "KnownEndpoints", "EndpointSettings", "MessageRedirects" })); + Assert.That(rowsAtFirstWrite!.Where(c => c.CategoryId != "KnownEndpoints").Select(c => c.State), Is.All.EqualTo(MigrationCategoryState.NotStarted)); + } + } + + [Test] + public async Task A_halted_category_is_left_untouched_and_the_next_category_still_runs() + { + var firstStarted = new DateTime(2026, 1, 1, 0, 0, 0, DateTimeKind.Utc); + var source = new InMemoryMigrationSource(); + var checkpointStore = new InMemoryMigrationCheckpointStore(); + var halted = await checkpointStore.Upsert(new MigrationCheckpoint("KnownEndpoints", MigrationCategoryState.Halted, "KnownEndpoints-1", 1, 0, null, null, firstStarted, firstStarted, firstStarted, "Halted: earlier run")); + var target = new InMemoryMigrationTarget(checkpointStore); + var engine = new MigrationEngine(source, target, checkpointStore, new FakeTimeProvider(), + new MigrationEngineOptions(TimeSpan.Zero, 5, 100, []), NullLogger.Instance); + source.Seed("KnownEndpoints", Row("KnownEndpoints-1"), Row("KnownEndpoints-2")); + // MessageRedirects does not follow KnownEndpoints, so a Failed KnownEndpoints cannot hold it back. + source.Seed("MessageRedirects", Row("MessageRedirects-1")); + + var results = await engine.RunCategories([MigrationCategoryRegistry.Find("KnownEndpoints")!, MigrationCategoryRegistry.Find("MessageRedirects")!]); + + using (Assert.EnterMultipleScope()) + { + Assert.That(target.RowsHandedToWrite("KnownEndpoints"), Is.Empty, "a Failed category waits for the operator"); + Assert.That(results[0], Is.EqualTo(halted), "the halted category's row was changed"); + Assert.That(await checkpointStore.Read("KnownEndpoints"), Is.EqualTo(halted)); + Assert.That(results[1].State, Is.EqualTo(MigrationCategoryState.Complete)); + Assert.That(target.WrittenRows("MessageRedirects").Select(row => row.SourceId), Is.EqualTo(new[] { "MessageRedirects-1" })); + } + } + + [Test] + public async Task A_finished_category_is_not_copied_again() + { + var source = new InMemoryMigrationSource(); + var checkpointStore = new InMemoryMigrationCheckpointStore(); + await checkpointStore.Upsert(new MigrationCheckpoint("KnownEndpoints", MigrationCategoryState.Complete, "KnownEndpoints-1", 1, 0, 1, null, DateTime.UtcNow, DateTime.UtcNow, DateTime.UtcNow, null)); + var target = new InMemoryMigrationTarget(checkpointStore); + var engine = new MigrationEngine(source, target, checkpointStore, new FakeTimeProvider(), + new MigrationEngineOptions(TimeSpan.Zero, 5, 100, []), NullLogger.Instance); + source.Seed("KnownEndpoints", Row("KnownEndpoints-1")); + + var results = await engine.RunCategories([MigrationCategoryRegistry.Find("KnownEndpoints")!, MigrationCategoryRegistry.Find("EndpointSettings")!]); + + using (Assert.EnterMultipleScope()) + { + Assert.That(target.RowsHandedToWrite("KnownEndpoints"), Is.Empty); + Assert.That(results[0].State, Is.EqualTo(MigrationCategoryState.Complete)); + } + } + [Test] public async Task Running_no_categories_copies_nothing_and_reports_nothing() { diff --git a/src/ServiceControl.UnitTests/Migration/MigrationEngineSkipReasonTests.cs b/src/ServiceControl.UnitTests/Migration/MigrationEngineSkipReasonTests.cs index 2e07501603..374a06de68 100644 --- a/src/ServiceControl.UnitTests/Migration/MigrationEngineSkipReasonTests.cs +++ b/src/ServiceControl.UnitTests/Migration/MigrationEngineSkipReasonTests.cs @@ -25,10 +25,10 @@ public async Task Reasons_the_target_reports_add_up_across_batches_on_the_checkp source.Seed(category.Id, Row("a"), Row("b"), Row("c"), Row("d")); var checkpointStore = new InMemoryMigrationCheckpointStore(); var target = new InMemoryMigrationTarget(checkpointStore) { DefaultBatchSize = 3 }; - // One reason exists, so this pins the total rather than the split between reasons. - target.RejectKey("a", MigrationSkipReason.BodyUnreadable); - target.RejectKey("b", MigrationSkipReason.BodyUnreadable); - target.RejectKey("d", MigrationSkipReason.BodyUnreadable); + // Two reasons over two batches: a and b land in the first, d in the second, so the saved map has to merge both. + target.RejectKey("a", MigrationSkipReason.RequiredValueMissing); + target.RejectKey("b", MigrationSkipReason.RequiredValueMissing); + target.RejectKey("d", MigrationSkipReason.PastRetention); var options = new MigrationEngineOptions(TimeSpan.Zero, 5, 100, []); var engine = new MigrationEngine(source, target, checkpointStore, new FakeTimeProvider(), options, NullLogger.Instance); @@ -37,8 +37,8 @@ public async Task Reasons_the_target_reports_add_up_across_batches_on_the_checkp using (Assert.EnterMultipleScope()) { Assert.That(checkpoint.SkippedCount, Is.EqualTo(3)); - Assert.That(checkpoint.SkipReasons, Is.EquivalentTo(new Dictionary { [MigrationSkipReason.BodyUnreadable] = 3 })); - Assert.That((await checkpointStore.Read(category.Id))!.SkipReasons, Is.EquivalentTo(new Dictionary { [MigrationSkipReason.BodyUnreadable] = 3 })); + Assert.That(checkpoint.SkipReasons, Is.EquivalentTo(new Dictionary { [MigrationSkipReason.RequiredValueMissing] = 2, [MigrationSkipReason.PastRetention] = 1 })); + Assert.That((await checkpointStore.Read(category.Id))!.SkipReasons, Is.EquivalentTo(new Dictionary { [MigrationSkipReason.RequiredValueMissing] = 2, [MigrationSkipReason.PastRetention] = 1 })); } } @@ -71,7 +71,7 @@ public async Task A_batch_that_loses_rows_two_different_ways_records_both_reason source.FailBodyReads("msg-1", MigrationEngine.MaxBodyReadAttempts, new TimeoutException("body store unreachable")); var checkpointStore = new InMemoryMigrationCheckpointStore(); var target = new InMemoryMigrationTarget(checkpointStore) { DefaultBatchSize = 3 }; - target.RejectKey("msg-2", MigrationSkipReason.PastRetention, benign: true); + target.RejectKey("msg-2", MigrationSkipReason.PastRetention); var options = new MigrationEngineOptions(TimeSpan.Zero, 5, 100, []) { BodyRetryBackoff = TimeSpan.Zero }; var engine = new MigrationEngine(source, target, checkpointStore, new FakeTimeProvider(), options, NullLogger.Instance); @@ -112,42 +112,82 @@ public async Task A_target_whose_skip_reasons_do_not_add_up_to_its_skips_halts_t } [Test] - public async Task A_target_counting_more_benign_skips_than_skipped_rows_halts_the_category() + public async Task A_target_whose_reported_counts_disagree_with_the_checkpoint_it_committed_never_finishes_the_category() { var category = MigrationCategoryRegistry.Find(MigrationCategoryIds.KnownEndpoints)!; var source = new InMemoryMigrationSource(); - source.Seed(category.Id, Row("a")); + source.Seed(category.Id, [.. Enumerable.Range(1, 300).Select(i => Row($"row-{i}"))]); var checkpointStore = new InMemoryMigrationCheckpointStore(); - var target = new OverCountedBenignTarget(checkpointStore); + var target = new MiscountedCopyTarget(checkpointStore); var options = new MigrationEngineOptions(TimeSpan.Zero, 5, 100, []); var engine = new MigrationEngine(source, target, checkpointStore, new FakeTimeProvider(), options, NullLogger.Instance); - var checkpoint = await engine.RunCategoryAsync(category); + var settled = await engine.RunCategoryAsync(category); + + var stored = await checkpointStore.Read(category.Id); using (Assert.EnterMultipleScope()) { - Assert.That(checkpoint.State, Is.EqualTo(MigrationCategoryState.Halted)); - Assert.That(checkpoint.LastError, Does.Contain("2 benign skips").And.Contain("out of 1 skipped")); + Assert.That(settled.State, Is.EqualTo(MigrationCategoryState.Halted)); + Assert.That(stored!.LastError, Does.Contain("but the checkpoint it committed moved by"), "the halt has to name which two accounts disagreed, or the operator goes looking for the wrong problem"); + Assert.That(stored.State.IsFinished(), Is.False, "the halt threshold never saw the 150 rows the target dropped, so the category settled finished and the host opened on half a category"); + } + } + + [Test] + public async Task An_optional_category_whose_reported_counts_disagree_with_the_checkpoint_it_committed_settles_halted() + { + // An optional category stays copying after an exception, and this check must not become one of those. + var category = MigrationCategoryRegistry.Find(MigrationCategoryIds.EventLog)!; + var source = new InMemoryMigrationSource(); + source.Seed(category.Id, [.. Enumerable.Range(1, 300).Select(i => Row($"row-{i}"))]); + var checkpointStore = new InMemoryMigrationCheckpointStore(); + var target = new MiscountedCopyTarget(checkpointStore); + var options = new MigrationEngineOptions(TimeSpan.Zero, 5, 100, []); + var engine = new MigrationEngine(source, target, checkpointStore, new FakeTimeProvider(), options, NullLogger.Instance); + + var settled = await engine.RunCategoryAsync(category); + + var stored = await checkpointStore.Read(category.Id); + + using (Assert.EnterMultipleScope()) + { + Assert.That(category.Kind, Is.EqualTo(MigrationCategoryKind.Optional)); + Assert.That(settled.State, Is.EqualTo(MigrationCategoryState.Halted)); + Assert.That(stored!.State, Is.EqualTo(MigrationCategoryState.Halted)); + Assert.That(stored.LastError, Does.Contain("reported copying 10").And.Contain("but the checkpoint it committed moved by")); + Assert.That(stored.LastError, Does.Not.Contain("--migration-retry").And.Not.Contain("--migration-abandon")); } } - sealed class OverCountedBenignTarget(IMigrationCheckpointStore checkpointStore) : IMigrationTarget + // Commits half of every batch as skipped and reports the whole batch copied. The saved counts still add up + // to the source total, so nothing later in the run can notice, and the halt threshold sees a clean copy. + sealed class MiscountedCopyTarget(IMigrationCheckpointStore checkpointStore) : IMigrationTarget { - public int BatchSizeFor(MigrationCategory category) => 10; + public Task Open(CancellationToken cancellationToken = default) => Task.CompletedTask; + + public Task BatchSizeFor(MigrationCategory category, CancellationToken cancellationToken = default) => Task.FromResult(10); public async Task Write(MigrationCategory category, MigrationBatch batch, MigrationCheckpoint checkpointToExtend, CancellationToken cancellationToken = default) { - var reasons = new Dictionary { [MigrationSkipReason.PastRetention] = batch.Rows.Count }; - var saved = await checkpointStore.Upsert(checkpointToExtend.Extend(0, batch.Rows.Count, 0, reasons), cancellationToken); - return new MigrationWriteResult(saved, 0, batch.Rows.Count, [], 0, reasons, BenignSkipped: batch.Rows.Count + 1); + var skipped = batch.Rows.Count / 2; + var reasons = new Dictionary { [MigrationSkipReason.RequiredValueMissing] = skipped }; + var saved = await checkpointStore.Upsert(checkpointToExtend.Extend(batch.Rows.Count - skipped, skipped, 0, reasons), cancellationToken); + return new MigrationWriteResult(saved, batch.Rows.Count, 0, []); } public Task Count(MigrationCategory category, CancellationToken cancellationToken = default) => Task.FromResult(0L); + + public IReadOnlyCollection SupportedCategoryIds => [.. MigrationCategoryRegistry.All.Select(category => category.Id)]; + + public IReadOnlyDictionary DocumentTypes => MigrationCategoryRegistry.All.ToDictionary(category => category.Id, _ => typeof(object)); } sealed class UnexplainedSkipTarget(IMigrationCheckpointStore checkpointStore) : IMigrationTarget { - public int BatchSizeFor(MigrationCategory category) => 10; + public Task Open(CancellationToken cancellationToken = default) => Task.CompletedTask; + + public Task BatchSizeFor(MigrationCategory category, CancellationToken cancellationToken = default) => Task.FromResult(10); public async Task Write(MigrationCategory category, MigrationBatch batch, MigrationCheckpoint checkpointToExtend, CancellationToken cancellationToken = default) { @@ -156,5 +196,9 @@ public async Task Write(MigrationCategory category, Migrat } public Task Count(MigrationCategory category, CancellationToken cancellationToken = default) => Task.FromResult(0L); + + public IReadOnlyCollection SupportedCategoryIds => [.. MigrationCategoryRegistry.All.Select(category => category.Id)]; + + public IReadOnlyDictionary DocumentTypes => MigrationCategoryRegistry.All.ToDictionary(category => category.Id, _ => typeof(object)); } } diff --git a/src/ServiceControl.UnitTests/Migration/MigrationIsReleasedCheckTests.cs b/src/ServiceControl.UnitTests/Migration/MigrationIsReleasedCheckTests.cs new file mode 100644 index 0000000000..638efaaf1b --- /dev/null +++ b/src/ServiceControl.UnitTests/Migration/MigrationIsReleasedCheckTests.cs @@ -0,0 +1,40 @@ +namespace ServiceControl.UnitTests.Migration; + +using System; +using Microsoft.Extensions.DependencyInjection; +using NUnit.Framework; +using ServiceBus.Management.Infrastructure.Settings; +using ServiceControl.Migration; +using ServiceControl.Migration.Checks; +using ServiceControl.Persistence.DataMigration; + +[TestFixture] +class MigrationIsReleasedCheckTests +{ + [Test] + public void A_build_without_the_whole_migration_refuses_MigrationEnabled() + { + var exception = Assert.ThrowsAsync(() => new MigrationIsReleasedCheck().Run()); + + Assert.That(exception.Message, Does.Contain("does not yet carry the whole migration").And.Contain(MigrationSettings.EnabledKey), + "the customer has to learn that nothing will be copied, and which setting to turn back off"); + } + + // A RavenDB target fails the pair check, so this only passes while the release check runs before it. + [Test] + public void An_unreleased_build_refuses_before_any_other_check() + { + var settings = new Settings(transportType: "LearningTransport", persisterType: "RavenDB", errorRetentionPeriod: TimeSpan.FromDays(10)); + using var services = new ServiceCollection().BuildServiceProvider(); + + var exception = Assert.ThrowsAsync(() => MigrationStartup.RunRequiredCopy(services, settings)); + + Assert.That(exception.Message, Does.Contain("this build carries the whole migration").And.Not.Contain("supported pair")); + } + + [Test] + public void The_test_marker_lets_a_copy_run_on_an_unreleased_build() + { + Assert.DoesNotThrowAsync(() => new MigrationIsReleasedCheck(new AllowUnreleasedMigration()).Run()); + } +} diff --git a/src/ServiceControl.UnitTests/Migration/MigrationPairIsSupportedCheckTests.cs b/src/ServiceControl.UnitTests/Migration/MigrationPairIsSupportedCheckTests.cs new file mode 100644 index 0000000000..c25bf44d0c --- /dev/null +++ b/src/ServiceControl.UnitTests/Migration/MigrationPairIsSupportedCheckTests.cs @@ -0,0 +1,35 @@ +namespace ServiceControl.UnitTests.Migration; + +using System; +using System.Linq; +using NUnit.Framework; +using ServiceBus.Management.Infrastructure.Settings; +using ServiceControl.Migration.Checks; +using ServiceControl.Persistence; + +[TestFixture] +class MigrationPairIsSupportedCheckTests +{ + [OneTimeSetUp] + public static void TheTreeIsBuilt() => + Assert.That( + PersistenceFactory.SqlPersistenceNames.Where(name => PersistenceManifestLibrary.Find(name) is null), + Is.Empty, + "These persistence manifests did not resolve, so this fixture cannot tell a supported pair from an unsupported one. Build the tree first with: dotnet build src --configuration Release -graph"); + + [TestCase("SQLServer")] + [TestCase("PostgreSQL")] + public void RavenDB_to_a_SQL_persister_is_supported(string target) => + Assert.DoesNotThrowAsync(() => Check(target).Run()); + + [Test] + public void A_RavenDB_target_is_refused_naming_the_target_setting() + { + var exception = Assert.ThrowsAsync(() => Check("RavenDB").Run()); + + Assert.That(exception.Message, Does.Contain("ServiceControl/PersistenceType")); + } + + static MigrationPairIsSupportedCheck Check(string target) => + new(new Settings(transportType: "LearningTransport", persisterType: target, errorRetentionPeriod: TimeSpan.FromDays(10))); +} diff --git a/src/ServiceControl.UnitTests/Migration/MigrationStartupCheckRunnerTests.cs b/src/ServiceControl.UnitTests/Migration/MigrationStartupCheckRunnerTests.cs new file mode 100644 index 0000000000..2767013bb3 --- /dev/null +++ b/src/ServiceControl.UnitTests/Migration/MigrationStartupCheckRunnerTests.cs @@ -0,0 +1,75 @@ +#nullable enable +namespace ServiceControl.UnitTests.Migration; + +using System; +using System.Collections.Generic; +using System.Threading; +using System.Threading.Tasks; +using NUnit.Framework; +using ServiceControl.Migration; +using ServiceControl.Persistence.DataMigration; + +[TestFixture] +class MigrationStartupCheckRunnerTests +{ + [Test] + public async Task Checks_run_in_the_order_they_were_given() + { + var ran = new List(); + + await MigrationStartupCheckRunner.Run( + [ + Check("the first", _ => ran.Add("first")), + Check("the second", _ => ran.Add("second")), + Check("the third", _ => ran.Add("third")) + ]); + + Assert.That(ran, Is.EqualTo(new[] { "first", "second", "third" }), "the databases are opened by checks that sit behind the ones refusing the configuration outright"); + } + + [Test] + public void The_first_check_to_refuse_stops_the_ones_behind_it() + { + var ran = new List(); + var failure = new InvalidOperationException("ServiceControl/RetryHistoryDepth is 0"); + + var exception = Assert.ThrowsAsync(() => MigrationStartupCheckRunner.Run( + [ + Check("the first", _ => ran.Add("first")), + Check("the migration target is ready", _ => throw failure), + Check("the last", _ => ran.Add("last")) + ])); + + using (Assert.EnterMultipleScope()) + { + Assert.That(ran, Is.EqualTo(new[] { "first" }), "a check behind a refusal would run against a configuration already known to be wrong, and the two that open a database are at the end of the list"); + Assert.That(exception!.Message, Does.Contain("the migration target is ready").And.Contain(failure.Message), "the operator needs to know which check refused and why"); + Assert.That(exception.InnerException, Is.SameAs(failure), "the original carries the stack and the detail the summary leaves out"); + } + } + + [Test] + public void A_host_being_stopped_is_not_a_check_refusing() + { + using var stopping = new CancellationTokenSource(); + stopping.Cancel(); + + var exception = Assert.ThrowsAsync(() => MigrationStartupCheckRunner.Run( + [Check("the first", token => token.ThrowIfCancellationRequested())], stopping.Token)); + + Assert.That(exception!.Message, Does.Not.Contain("Migration startup check"), "a stop dressed up as a refusal sends the operator after a setting that was never wrong"); + } + + static IMigrationStartupCheck Check(string name, Action run) => new FakeCheck(name, run); + + sealed class FakeCheck(string name, Action run) : IMigrationStartupCheck + { + public string Name => name; + + public Task Run(CancellationToken cancellationToken = default) + { + run(cancellationToken); + return Task.CompletedTask; + } + } +} diff --git a/src/ServiceControl.UnitTests/Migration/OptionalCategoryWindowsAreValidCheckTests.cs b/src/ServiceControl.UnitTests/Migration/OptionalCategoryWindowsAreValidCheckTests.cs new file mode 100644 index 0000000000..6f4393ab40 --- /dev/null +++ b/src/ServiceControl.UnitTests/Migration/OptionalCategoryWindowsAreValidCheckTests.cs @@ -0,0 +1,51 @@ +namespace ServiceControl.UnitTests.Migration; + +using System; +using System.Threading.Tasks; +using NUnit.Framework; +using ServiceControl.Migration.Checks; +using ServiceControl.Persistence.DataMigration; + +[TestFixture] +[NonParallelizable] +class OptionalCategoryWindowsAreValidCheckTests +{ + static readonly TimeSpan EventRetention = TimeSpan.FromDays(3); + static readonly TimeSpan ErrorRetention = TimeSpan.FromDays(11); + + [TearDown] + public void ClearWindows() + { + Environment.SetEnvironmentVariable("SERVICECONTROL_MIGRATION_EVENTLOGWINDOW", null); + Environment.SetEnvironmentVariable("SERVICECONTROL_MIGRATION_ARCHIVEDANDRESOLVEDFAILEDMESSAGESWINDOW", null); + } + + [Test] + public async Task Unset_windows_pass_and_default_to_the_retention_periods_it_was_given() + { + var check = new OptionalCategoryWindowsAreValidCheck(EventRetention, ErrorRetention); + + await check.Run(); + + using (Assert.EnterMultipleScope()) + { + Assert.That(check.Options.EventLogWindow, Is.EqualTo(EventRetention)); + Assert.That(check.Options.ArchivedAndResolvedFailedMessagesWindow, Is.EqualTo(ErrorRetention)); + } + } + + [Test] + public void A_window_that_is_not_a_time_span_is_refused_and_named() + { + Environment.SetEnvironmentVariable("SERVICECONTROL_MIGRATION_ARCHIVEDANDRESOLVEDFAILEDMESSAGESWINDOW", "a week"); + + var exception = Assert.ThrowsAsync(async () => + await new OptionalCategoryWindowsAreValidCheck(EventRetention, ErrorRetention).Run()); + + using (Assert.EnterMultipleScope()) + { + Assert.That(exception.Message, Does.Contain("a week")); + Assert.That(exception.Message, Does.Contain(MigrationSettings.ArchivedAndResolvedFailedMessagesWindowKey)); + } + } +} diff --git a/src/ServiceControl.UnitTests/Migration/RecordSourceOutageTests.cs b/src/ServiceControl.UnitTests/Migration/RecordSourceOutageTests.cs new file mode 100644 index 0000000000..cc598257ab --- /dev/null +++ b/src/ServiceControl.UnitTests/Migration/RecordSourceOutageTests.cs @@ -0,0 +1,80 @@ +#nullable enable +namespace ServiceControl.UnitTests.Migration; + +using System; +using System.Threading.Tasks; +using NUnit.Framework; +using ServiceControl.Migration; +using ServiceControl.Persistence.DataMigration; +using ServiceControl.UnitTests.Migration.Fakes; + +[TestFixture] +class RecordSourceOutageTests +{ + static readonly MigrationCategory EventLog = MigrationCategoryRegistry.Find(MigrationCategoryIds.EventLog)!; + static readonly MigrationCategory ArchivedAndResolvedFailedMessages = MigrationCategoryRegistry.Find(MigrationCategoryIds.ArchivedAndResolvedFailedMessages)!; + + readonly InMemoryMigrationCheckpointStore checkpointStore = new(); + + [Test] + public async Task A_category_still_copying_keeps_its_state_and_gains_the_error() + { + await checkpointStore.Upsert(new MigrationCheckpoint(MigrationCategoryIds.EventLog, MigrationCategoryState.InProgress, "EventLogItem/40", 40, 0, null, null, null, null, null, null)); + + await RecordOutage(); + + var row = await checkpointStore.Read(MigrationCategoryIds.EventLog); + + using (Assert.EnterMultipleScope()) + { + Assert.That(row!.State, Is.EqualTo(MigrationCategoryState.InProgress)); + Assert.That(row.Cursor, Is.EqualTo("EventLogItem/40")); + Assert.That(row.CopiedCount, Is.EqualTo(40)); + Assert.That(row.SkippedCount, Is.Zero); + Assert.That(row.LastError, Does.Contain("RavenDB is unreachable").And.Contain("next start")); + } + } + + [Test] + public async Task A_category_with_no_row_gets_a_not_started_row_carrying_the_error() + { + await RecordOutage(); + + var row = await checkpointStore.Read(MigrationCategoryIds.ArchivedAndResolvedFailedMessages); + + Assert.That(row, Is.Not.Null, "status and verify read only the rows that exist, so a category with no row would show no error"); + + using (Assert.EnterMultipleScope()) + { + Assert.That(row!.State, Is.EqualTo(MigrationCategoryState.NotStarted)); + Assert.That(row.CopiedCount, Is.Zero); + Assert.That(row.SkippedCount, Is.Zero); + Assert.That(row.AlreadyPresentCount, Is.Zero); + Assert.That(row.StartedAt, Is.Null, "a category that never started must not look as if it had"); + Assert.That(row.LastError, Does.Contain("RavenDB is unreachable").And.Contain("next start")); + } + } + + [TestCase(MigrationCategoryState.Complete)] + [TestCase(MigrationCategoryState.Halted)] + [TestCase(MigrationCategoryState.CompleteWithErrors)] + [TestCase(MigrationCategoryState.Abandoned)] + public async Task A_settled_category_is_left_alone(MigrationCategoryState state) + { + var before = await checkpointStore.Upsert(new MigrationCheckpoint(MigrationCategoryIds.EventLog, state, null, 0, 0, null, null, null, null, null, "earlier")); + + await RecordOutage(); + + var after = await checkpointStore.Read(MigrationCategoryIds.EventLog); + + using (Assert.EnterMultipleScope()) + { + Assert.That(after!.Version, Is.EqualTo(before.Version), "a settled row is never saved again"); + Assert.That(after.State, Is.EqualTo(state)); + Assert.That(after.LastError, Is.EqualTo("earlier")); + } + } + + Task RecordOutage() => + MigrationStartup.RecordSourceOutage(checkpointStore, [EventLog, ArchivedAndResolvedFailedMessages], new InvalidOperationException("RavenDB is unreachable")); +} diff --git a/src/ServiceControl.UnitTests/Migration/RequiredCopyGateTests.cs b/src/ServiceControl.UnitTests/Migration/RequiredCopyGateTests.cs new file mode 100644 index 0000000000..2077493bff --- /dev/null +++ b/src/ServiceControl.UnitTests/Migration/RequiredCopyGateTests.cs @@ -0,0 +1,395 @@ +#nullable enable +namespace ServiceControl.UnitTests.Migration; + +using System; +using System.Collections.Generic; +using System.Linq; +using Microsoft.Extensions.Logging; +using NUnit.Framework; +using ServiceBus.Management.Infrastructure.Settings; +using ServiceControl.Migration; +using ServiceControl.Persistence.DataMigration; +using ServiceControl.UnitTests.Migration.Fakes; + +[TestFixture] +class RequiredCopyGateTests +{ + [TestCase(MigrationCategoryState.Complete)] + public void A_finished_required_category_lets_the_host_open(MigrationCategoryState state) => + Assert.DoesNotThrow(() => MigrationStartup.RefuseIfAnyCategoryDidNotComplete([EndpointSettings], [Checkpoint(state)], NewSettings())); + + [TestCase(MigrationCategoryState.Halted)] + [TestCase(MigrationCategoryState.CompleteWithErrors)] + [TestCase(MigrationCategoryState.InProgress)] + [TestCase(MigrationCategoryState.NotStarted)] + // Blocked happens because EndpointSettings must follow KnownEndpoints. + [TestCase(MigrationCategoryState.Blocked)] + public void A_required_category_that_is_not_finished_keeps_the_host_closed(MigrationCategoryState state) + { + var exception = Assert.Throws(() => MigrationStartup.RefuseIfAnyCategoryDidNotComplete([EndpointSettings], [Checkpoint(state)], NewSettings())); + + Assert.That(exception.Message, Does.Contain(MigrationCategoryIds.EndpointSettings).And.Contain(state.ToString())); + } + + [TestCase(MigrationCategoryState.InProgress)] + [TestCase(MigrationCategoryState.Halted)] + public void An_unfinished_row_under_an_id_this_build_does_not_know_keeps_the_host_closed(MigrationCategoryState state) + { + var unknown = Checkpoint(state, UnknownCategoryId); + + var outside = MigrationStartup.UnfinishedRowsOutsideTheCopy([unknown, Checkpoint(MigrationCategoryState.Complete)], [EndpointSettings]); + var exception = Assert.Throws(() => MigrationStartup.RefuseIfAnyCategoryDidNotComplete([EndpointSettings], [Checkpoint(MigrationCategoryState.Complete), .. outside], NewSettings())); + + Assert.That(exception.Message, Does.Contain($"{UnknownCategoryId} is ").And.Contain(state.ToString()), + "an id this build does not know is required until it is known to be optional, as the ingestion gate already decides"); + } + + [Test] + public void An_unfinished_unknown_row_makes_the_copy_not_settled_and_the_refusal_names_it() + { + var unknown = Checkpoint(MigrationCategoryState.InProgress, UnknownCategoryId); + var checkpoints = new[] { Checkpoint(MigrationCategoryState.Complete), unknown }; + var outside = MigrationStartup.UnfinishedRowsOutsideTheCopy(checkpoints, [EndpointSettings]); + + var exception = Assert.Throws(() => MigrationStartup.RefuseIfAnyCategoryDidNotComplete([EndpointSettings], [Checkpoint(MigrationCategoryState.Complete)], outside, NewSettings())); + + Assert.Multiple(() => + { + Assert.That(MigrationStartup.RequiredCopyIsSettled(checkpoints, [EndpointSettings]), Is.False); + Assert.That(exception.Message, Does.Contain($"{UnknownCategoryId} is InProgress")); + }); + } + + [Test] + public void A_finished_unknown_row_leaves_the_copy_settled_and_the_refusal_silent() + { + var checkpoints = new[] { Checkpoint(MigrationCategoryState.Complete), Checkpoint(MigrationCategoryState.Abandoned, UnknownCategoryId) }; + var outside = MigrationStartup.UnfinishedRowsOutsideTheCopy(checkpoints, [EndpointSettings]); + + Assert.Multiple(() => + { + Assert.That(MigrationStartup.RequiredCopyIsSettled(checkpoints, [EndpointSettings]), Is.True); + Assert.DoesNotThrow(() => MigrationStartup.RefuseIfAnyCategoryDidNotComplete([EndpointSettings], [Checkpoint(MigrationCategoryState.Complete)], outside, NewSettings())); + }); + } + + [TestCase(MigrationCategoryState.Complete)] + [TestCase(MigrationCategoryState.Abandoned)] + public void A_finished_row_under_an_id_this_build_does_not_know_holds_nothing(MigrationCategoryState state) => + Assert.That(MigrationStartup.UnfinishedRowsOutsideTheCopy([Checkpoint(state, UnknownCategoryId)], [EndpointSettings]), Is.Empty); + + [Test] + public void An_unfinished_row_known_to_be_optional_holds_nothing() => + Assert.That(MigrationStartup.UnfinishedRowsOutsideTheCopy([Checkpoint(MigrationCategoryState.InProgress, MigrationCategoryIds.EventLog)], [EndpointSettings]), Is.Empty); + + [Test] + public void A_category_that_reported_no_checkpoint_at_all_keeps_the_host_closed() + { + var exception = Assert.Throws(() => MigrationStartup.RefuseIfAnyCategoryDidNotComplete([EndpointSettings], [], NewSettings())); + + Assert.That(exception.Message, Does.Contain(MigrationCategoryIds.EndpointSettings), + "an empty result reads as success unless the gate checks what it asked for against what came back"); + } + + [Test] + public void Of_two_attempted_categories_the_one_that_reported_nothing_is_the_one_named() + { + var exception = Assert.Throws(() => MigrationStartup.RefuseIfAnyCategoryDidNotComplete( + [KnownEndpoints, EndpointSettings], + [Checkpoint(MigrationCategoryState.Complete, MigrationCategoryIds.KnownEndpoints)], + NewSettings())); + + Assert.That(exception.Message, Does.Contain($"{MigrationCategoryIds.EndpointSettings} reported no checkpoint at all").And.Not.Contain(MigrationCategoryIds.KnownEndpoints), + "comparing counts instead of ids would let a two-category run through with one category's fate unknown"); + } + + [Test] + public void The_refusal_reads_to_its_end_as_a_sentence_and_says_how_to_get_back() + { + var exception = Assert.Throws(() => MigrationStartup.RefuseIfAnyCategoryDidNotComplete([EndpointSettings], [Checkpoint(MigrationCategoryState.Halted)], NewSettings())); + + Assert.Multiple(() => + { + Assert.That(exception.Message, Does.Contain($"--migration-retry {MigrationCategoryIds.EndpointSettings}").And.Contain($"--migration-abandon {MigrationCategoryIds.EndpointSettings}"), + "a Failed category waits for the operator, so a refusal naming neither command leaves them restarting for ever"); + Assert.That(exception.Message, Does.Not.Contain("the copy resumes from its last committed batch"), "a restart no longer copies a Failed category"); + Assert.That(exception.Message, Does.Contain("both with ServiceControl stopped. Nothing has opened on SQLServer yet") + .And.EndWith("returns the instance to RavenDB with no loss."), + "the operator reads the whole message, and the rollback instructions are at the end of it"); + }); + } + + [Test] + public void Every_failed_required_category_is_named_with_its_skips_and_both_commands() + { + var withErrors = Checkpoint(MigrationCategoryState.CompleteWithErrors, MigrationCategoryIds.UnresolvedAndRetryIssuedFailedMessages) with + { + CopiedCount = 40, + SkippedCount = 5, + SkipReasons = new Dictionary + { + [MigrationSkipReason.RequiredValueMissing] = 2, + [MigrationSkipReason.BodyUnreadable] = 3 + } + }; + var halted = Checkpoint(MigrationCategoryState.Halted, MigrationCategoryIds.MessageRedirects) with { LastError = "TimeoutException at cursor r-7: the target did not answer" }; + + var exception = Assert.Throws(() => MigrationStartup.RefuseIfAnyCategoryDidNotComplete( + [MessageRedirects, UnresolvedAndRetryIssuedFailedMessages], [halted, withErrors], NewSettings())); + + using (Assert.EnterMultipleScope()) + { + foreach (var categoryId in new[] { MigrationCategoryIds.MessageRedirects, MigrationCategoryIds.UnresolvedAndRetryIssuedFailedMessages }) + { + Assert.That(exception.Message, Does.Contain($"--migration-retry {categoryId} once the cause is fixed").And.Contain($"--migration-abandon {categoryId} to keep what was copied"), + $"one start gives every required category its go, so the refusal has to name every Failed one, {categoryId} included"); + } + + Assert.That(exception.Message, Does.Contain($"{MigrationCategoryIds.MessageRedirects} is Failed (Halted)").And.Contain("the target did not answer")); + Assert.That(exception.Message, Does.Contain($"{MigrationCategoryIds.UnresolvedAndRetryIssuedFailedMessages} is Failed (CompleteWithErrors) after copying 40 and skipping 5")); + Assert.That(exception.Message, Does.Contain("RequiredValueMissing 2, no retry can fix"), "a row no retry can store is a reason to abandon, and the operator has to be told which"); + Assert.That(exception.Message, Does.Contain("BodyUnreadable 3").And.Not.Contain("BodyUnreadable 3, no retry can fix"), "an unreadable body may read on the next try"); + } + } + + [Test] + public void A_category_waiting_behind_a_failed_one_is_named_as_waiting() + { + var halted = Checkpoint(MigrationCategoryState.Halted, MigrationCategoryIds.KnownEndpoints); + var blocked = Checkpoint(MigrationCategoryState.Blocked) with { LastError = "Blocked: EndpointSettings must follow KnownEndpoints, which is Halted" }; + + var exception = Assert.Throws(() => MigrationStartup.RefuseIfAnyCategoryDidNotComplete([KnownEndpoints, EndpointSettings], [halted, blocked], NewSettings())); + + using (Assert.EnterMultipleScope()) + { + Assert.That(exception.Message, Does.Contain($"{MigrationCategoryIds.EndpointSettings} is Blocked").And.Contain("must follow KnownEndpoints")); + Assert.That(exception.Message, Does.Contain($"--migration-retry {MigrationCategoryIds.KnownEndpoints}")); + Assert.That(exception.Message, Does.Not.Contain($"--migration-retry {MigrationCategoryIds.EndpointSettings}").And.Not.Contain($"--migration-abandon {MigrationCategoryIds.EndpointSettings}"), + "a category that never got its go has nothing to retry; it runs once the one ahead of it is dealt with"); + } + } + + [Test] + public void A_stalled_category_is_refused_with_the_limit_its_row_records() + { + var stalled = Checkpoint(MigrationCategoryState.Halted) with { LastError = MigrationStartup.StallExplanation(MigrationCategoryIds.EndpointSettings) }; + + var exception = Assert.Throws(() => MigrationStartup.RefuseIfAnyCategoryDidNotComplete([EndpointSettings], [stalled], NewSettings())); + + Assert.That(exception.Message, Does.Contain($"committed nothing for {MigrationStartup.ClosedWindowProgress.StallLimit.TotalMinutes:0.#} minutes") + .And.Contain("looking for a setting to change; run --migration-retry EndpointSettings"), + "the stall is on the row, and the refusal is where the operator reads it"); + } + + // The last thing said before the cutover is one way, and the only place a skipped total appears at all. + [Test] + public void A_failed_category_that_skipped_rows_is_told_a_retry_re_reads_them() + { + var logger = new CapturingLogger(); + + MigrationStartup.ReportWhatTheCopyLeftBehind( + [Settled(MigrationCategoryState.CompleteWithErrors, copied: 7, skipped: 6, new Dictionary + { + [MigrationSkipReason.PastRetention] = 3, + [MigrationSkipReason.RequiredValueMissing] = 2, + [MigrationSkipReason.BodyUnreadable] = 1 + })], + logger, + RunStartedAt); + + var entry = logger.Entries.Single(); + + Assert.Multiple(() => + { + Assert.That(entry.Level, Is.EqualTo(LogLevel.Warning), "rows left behind are not an informational matter"); + Assert.That(entry.Message, Does.Contain("6 skipped"), "the total is the number the operator decides on"); + Assert.That(entry.Message, Does.Contain("PastRetention 3").And.Contain("RequiredValueMissing 2").And.Contain("BodyUnreadable 1"), "a total with no reasons cannot be acted on"); + Assert.That(entry.Message, Does.Contain("stay only in the source database").And.Contain($"--migration-retry {MigrationCategoryIds.EndpointSettings} re-reads them"), + "without this the operator either waits for a copy that is already over or never learns the rows can still come across"); + Assert.That(entry.Message, Does.Contain("No retry can fix the RequiredValueMissing ones"), "the retry brings the unreadable bodies across but leaves these behind again"); + }); + } + + [Test] + public void A_failed_category_whose_faults_no_retry_can_fix_is_pointed_at_abandon() + { + var logger = new CapturingLogger(); + + MigrationStartup.ReportWhatTheCopyLeftBehind( + [Settled(MigrationCategoryState.CompleteWithErrors, copied: 7, skipped: 5, new Dictionary + { + [MigrationSkipReason.PastRetention] = 3, + [MigrationSkipReason.RequiredValueMissing] = 2 + })], + logger, + RunStartedAt); + + var entry = logger.Entries.Single(); + + Assert.Multiple(() => + { + Assert.That(entry.Message, Does.Contain("no retry can fix them").And.Contain($"--migration-abandon {MigrationCategoryIds.EndpointSettings}"), + "a retry re-reads the whole category and ends Failed on the same rows"); + Assert.That(entry.Message, Does.Not.Contain("--migration-retry"), "harmless skips are not a reason to retry either"); + }); + } + + [Test] + public void A_category_that_failed_with_only_harmless_skips_is_still_offered_a_retry() + { + var logger = new CapturingLogger(); + + var stoppedByAnException = Settled(MigrationCategoryState.Halted, copied: 7, skipped: 3, new Dictionary { [MigrationSkipReason.PastRetention] = 3 }) + with + { LastError = "TimeoutException at cursor r-7: the target did not answer" }; + + MigrationStartup.ReportWhatTheCopyLeftBehind([stoppedByAnException], logger, RunStartedAt); + + Assert.That(logger.Entries.Single().Message, Does.Contain($"--migration-retry {MigrationCategoryIds.EndpointSettings} re-reads them").And.Not.Contain("--migration-abandon"), + "no fault skip stands in the way of a retry, so the exception's cause is what the operator fixes"); + } + + [Test] + public void A_done_category_with_harmless_skips_is_not_offered_a_retry() + { + var logger = new CapturingLogger(); + + MigrationStartup.ReportWhatTheCopyLeftBehind( + [Settled(MigrationCategoryState.Complete, copied: 7, skipped: 3, new Dictionary { [MigrationSkipReason.PastRetention] = 3 })], + logger, + RunStartedAt); + + var entry = logger.Entries.Single(); + + Assert.Multiple(() => + { + Assert.That(entry.Message, Does.Contain("3 skipped").And.Contain("PastRetention 3")); + Assert.That(entry.Message, Does.Contain("harmless").And.Contain("stay only in the source database")); + Assert.That(entry.Message, Does.Not.Contain("--migration-retry"), "a Done category cannot be retried, so offering it sends the operator to a command that refuses"); + }); + } + + [Test] + public void A_blocked_category_is_left_to_the_refusal_rather_than_reported_as_a_clean_copy() + { + var logger = new CapturingLogger(); + + MigrationStartup.ReportWhatTheCopyLeftBehind([Settled(MigrationCategoryState.Blocked, copied: 0, skipped: 0, null)], logger, RunStartedAt); + + Assert.That(logger.Entries, Is.Empty); + } + + // Every refused start reports it again, directly above the refusal that calls it Failed. + [Test] + public void A_failed_category_that_skipped_nothing_is_not_reported_as_a_clean_copy() + { + var logger = new CapturingLogger(); + + var stalled = Settled(MigrationCategoryState.Halted, copied: 10, skipped: 0, null) with { LastError = MigrationStartup.StallExplanation(MigrationCategoryIds.EndpointSettings) }; + + MigrationStartup.ReportWhatTheCopyLeftBehind([stalled], logger, RunStartedAt); + + var entry = logger.Entries.Single(); + + Assert.Multiple(() => + { + Assert.That(entry.Level, Is.EqualTo(LogLevel.Warning)); + Assert.That(entry.Message, Does.Contain($"{MigrationCategoryIds.EndpointSettings} is Failed (Halted) after 10 copied").And.Not.Contain("nothing skipped")); + }); + } + + [Test] + public void A_category_that_skipped_nothing_says_so_without_raising_a_warning() + { + var logger = new CapturingLogger(); + + MigrationStartup.ReportWhatTheCopyLeftBehind([Settled(MigrationCategoryState.Complete, copied: 7, skipped: 0, null)], logger, RunStartedAt); + + var entry = logger.Entries.Single(); + + Assert.Multiple(() => + { + Assert.That(entry.Level, Is.EqualTo(LogLevel.Information), "a clean copy warning about nothing trains the operator to ignore the warning that matters"); + Assert.That(entry.Message, Does.Contain("nothing skipped")); + }); + } + + // Restarting a migrated instance used to reprint the whole copy summary in the present tense, so an operator + // restarting to change a setting read "7 copied" and had no way to tell the copy had not run again. + [Test] + public void A_category_an_earlier_run_finished_is_not_reported_as_copied_again() + { + var logger = new CapturingLogger(); + + var alreadyDone = Settled(MigrationCategoryState.Complete, copied: 7, skipped: 0, null) + with + { SettledAt = RunStartedAt.AddMinutes(-5) }; + + MigrationStartup.ReportWhatTheCopyLeftBehind([alreadyDone], logger, RunStartedAt); + + var entry = logger.Entries.Single(); + + Assert.Multiple(() => + { + Assert.That(entry.Message, Does.Contain("already finished before this start"), "without this the line is indistinguishable from a copy that just ran"); + Assert.That(entry.Message, Does.Contain("This start copied nothing"), "the operator needs to know nothing was written to a target that is already serving"); + Assert.That(entry.Message, Does.Contain("7"), "the historical total is still worth stating, just not as this run's work"); + Assert.That(entry.Level, Is.EqualTo(LogLevel.Information)); + }); + } + + // A category this run actually settled must keep the ordinary wording, or the fix above would silence every report. + [Test] + public void A_category_this_run_finished_is_still_reported_as_copied() + { + var logger = new CapturingLogger(); + + var justDone = Settled(MigrationCategoryState.Complete, copied: 7, skipped: 0, null) + with + { SettledAt = RunStartedAt.AddSeconds(2) }; + + MigrationStartup.ReportWhatTheCopyLeftBehind([justDone], logger, RunStartedAt); + + Assert.That(logger.Entries.Single().Message, Does.Contain("7 copied").And.Not.Contain("already finished")); + } + + // A Failed row comes back untouched on every start, so every refused restart reports it again. + [Test] + public void A_category_an_earlier_run_left_failed_is_not_reported_as_finished() + { + var logger = new CapturingLogger(); + + var failedEarlier = Settled(MigrationCategoryState.Halted, copied: 7, skipped: 2, new Dictionary { [MigrationSkipReason.RequiredValueMissing] = 2 }) + with + { SettledAt = RunStartedAt.AddMinutes(-5) }; + + MigrationStartup.ReportWhatTheCopyLeftBehind([failedEarlier], logger, RunStartedAt); + + var entry = logger.Entries.Single(); + + Assert.Multiple(() => + { + Assert.That(entry.Message, Does.Not.Contain("already finished"), "the refusal right after this line calls the same category Failed"); + Assert.That(entry.Message, Does.Contain($"--migration-abandon {MigrationCategoryIds.EndpointSettings}").And.Not.Contain("--migration-retry"), + "the advice has to show on every refused start, not only the first, and no retry can store a row missing a required value"); + }); + } + + static readonly DateTime RunStartedAt = new(2026, 9, 20, 12, 0, 0, DateTimeKind.Utc); + + static MigrationCheckpoint Settled(MigrationCategoryState state, long copied, long skipped, IReadOnlyDictionary? skipReasons) => + new(MigrationCategoryIds.EndpointSettings, state, null, copied, skipped, null, skipReasons, null, null, null, null); + + static readonly MigrationCategory EndpointSettings = MigrationCategoryRegistry.Find(MigrationCategoryIds.EndpointSettings)!; + static readonly MigrationCategory KnownEndpoints = MigrationCategoryRegistry.Find(MigrationCategoryIds.KnownEndpoints)!; + static readonly MigrationCategory MessageRedirects = MigrationCategoryRegistry.Find(MigrationCategoryIds.MessageRedirects)!; + static readonly MigrationCategory UnresolvedAndRetryIssuedFailedMessages = MigrationCategoryRegistry.Find(MigrationCategoryIds.UnresolvedAndRetryIssuedFailedMessages)!; + + const string UnknownCategoryId = "SomeCategoryFromANewerBuild"; + + static MigrationCheckpoint Checkpoint(MigrationCategoryState state, string categoryId = MigrationCategoryIds.EndpointSettings) => + new(categoryId, state, null, 0, 0, null, null, null, null, null, null); + + static Settings NewSettings() => + new(transportType: "LearningTransport", persisterType: "SQLServer", errorRetentionPeriod: TimeSpan.FromDays(10)); +} diff --git a/src/ServiceControl.UnitTests/Migration/RetryHistoryDepthIsSafeCheckTests.cs b/src/ServiceControl.UnitTests/Migration/RetryHistoryDepthIsSafeCheckTests.cs new file mode 100644 index 0000000000..f3aa068120 --- /dev/null +++ b/src/ServiceControl.UnitTests/Migration/RetryHistoryDepthIsSafeCheckTests.cs @@ -0,0 +1,22 @@ +namespace ServiceControl.UnitTests.Migration; + +using System; +using NUnit.Framework; +using ServiceControl.Migration.Checks; + +[TestFixture] +class RetryHistoryDepthIsSafeCheckTests +{ + [Test] + public void Passes_at_the_default_depth() => + Assert.DoesNotThrowAsync(() => new RetryHistoryDepthIsSafeCheck(10).Run()); + + [TestCase(0)] + [TestCase(-1)] + public void Refuses_a_depth_that_empties_the_table(int depth) + { + var exception = Assert.ThrowsAsync(() => new RetryHistoryDepthIsSafeCheck(depth).Run()); + + Assert.That(exception.Message, Does.Contain("RetryHistoryDepth").And.Contain("HistoricRetryOperations")); + } +} diff --git a/src/ServiceControl.UnitTests/Migration/StoppedCopyExplanationTests.cs b/src/ServiceControl.UnitTests/Migration/StoppedCopyExplanationTests.cs new file mode 100644 index 0000000000..2f3b033642 --- /dev/null +++ b/src/ServiceControl.UnitTests/Migration/StoppedCopyExplanationTests.cs @@ -0,0 +1,194 @@ +#nullable enable +namespace ServiceControl.UnitTests.Migration; + +using System; +using System.Collections.Generic; +using System.Linq; +using System.Threading; +using System.Threading.Tasks; +using Microsoft.Extensions.Logging.Abstractions; +using Microsoft.Extensions.Time.Testing; +using NUnit.Framework; +using ServiceBus.Management.Infrastructure.Settings; +using ServiceControl.Migration; +using ServiceControl.Persistence.DataMigration; +using ServiceControl.UnitTests.Migration.Fakes; + +// The copy can stop three ways that are not the engine's to report: the stall watchdog stopping one category, +// the host shutting down, and a second instance writing to the same database. +[TestFixture] +class StoppedCopyExplanationTests +{ + static readonly TimeSpan PollInterval = MigrationStartup.ClosedWindowProgress.PollInterval; + static readonly TimeSpan StallLimit = MigrationStartup.ClosedWindowProgress.StallLimit; + + [Test] + public async Task A_stall_settles_the_stalled_category_failed_and_the_next_category_still_runs() + { + var clock = new TimerRecordingTimeProvider(); + var store = new CancellationHonouringCheckpointStore(); + var writing = new TaskCompletionSource(TaskCreationOptions.RunContinuationsAsynchronously); + var target = new InMemoryMigrationTarget(store) { HangOnCall = 1, BeforeWrite = _ => writing.TrySetResult() }; + var source = new InMemoryMigrationSource(); + source.Seed(MigrationCategoryIds.KnownEndpoints, Row("KnownEndpoints-1")); + // MessageRedirects does not follow KnownEndpoints, so a Failed KnownEndpoints cannot hold it back. + source.Seed(MigrationCategoryIds.MessageRedirects, Row("MessageRedirects-1")); + + var run = Task.Run(() => MigrationStartup.RunRequiredCategories( + Engine(source, target, store, clock), Categories(MigrationCategoryIds.KnownEndpoints, MigrationCategoryIds.MessageRedirects), store, clock, new CapturingLogger())); + + Assert.That(await clock.TimerCreated.WaitAsync(TimeSpan.FromSeconds(10)), Is.True, "the watchdog never started its timer, so advancing the clock would tick nothing"); + await writing.Task.WaitAsync(TimeSpan.FromSeconds(10)); + clock.Advance(StallLimit + PollInterval); + var results = await run.WaitAsync(TimeSpan.FromSeconds(10)); + + using (Assert.EnterMultipleScope()) + { + Assert.That(results.Select(checkpoint => checkpoint.CategoryId), Is.EqualTo(new[] { MigrationCategoryIds.KnownEndpoints, MigrationCategoryIds.MessageRedirects })); + Assert.That(results[0].State, Is.EqualTo(MigrationCategoryState.Halted), "a stalled category is Failed, so it waits for the operator rather than for the next start"); + Assert.That(results[0].SettledAt, Is.EqualTo(clock.GetUtcNow().UtcDateTime)); + Assert.That(results[0].LastError, Does.Contain($"committed nothing for {StallLimit.TotalMinutes:0.#} minutes").And.Not.Contain("--migration-"), + "the row says what happened, and the refusal is what names the commands"); + Assert.That(await store.Read(MigrationCategoryIds.KnownEndpoints), Is.EqualTo(results[0]), "the stall was reported but never saved, so the next start would copy it again"); + Assert.That(results[1].State, Is.EqualTo(MigrationCategoryState.Complete), "one stalled category ended the whole start"); + Assert.That(target.WrittenRows(MigrationCategoryIds.MessageRedirects).Select(row => row.SourceId), Is.EqualTo(new[] { "MessageRedirects-1" })); + } + } + + [Test] + public async Task A_host_stopping_mid_category_leaves_it_copying_and_the_stop_propagates() + { + var store = new InMemoryMigrationCheckpointStore(); + using var host = new CancellationTokenSource(); + var target = new InMemoryMigrationTarget(store) { DefaultBatchSize = 2, StopOnCall = (2, host) }; + var source = new InMemoryMigrationSource(); + source.Seed(MigrationCategoryIds.KnownEndpoints, Row("KnownEndpoints-1"), Row("KnownEndpoints-2"), Row("KnownEndpoints-3"), Row("KnownEndpoints-4")); + source.Seed(MigrationCategoryIds.MessageRedirects, Row("MessageRedirects-1")); + var clock = new FakeTimeProvider(); + + var run = MigrationStartup.RunRequiredCategories( + Engine(source, target, store, clock), Categories(MigrationCategoryIds.KnownEndpoints, MigrationCategoryIds.MessageRedirects), store, clock, new CapturingLogger(), host.Token); + + using (Assert.EnterMultipleScope()) + { + Assert.That(async () => await run, Throws.InstanceOf(), "a shutdown is not the category's fault, so it must not come back as a settled row"); + + var stopped = await store.Read(MigrationCategoryIds.KnownEndpoints); + + Assert.That(stopped!.State, Is.EqualTo(MigrationCategoryState.InProgress), "a shutdown settled Halted makes every restart refuse until the operator runs --migration-retry"); + Assert.That(stopped.Cursor, Is.EqualTo("KnownEndpoints-2"), "the next start resumes from the first batch's cursor"); + Assert.That(stopped.CopiedCount, Is.EqualTo(2)); + Assert.That(stopped.SettledAt, Is.Null); + Assert.That(stopped.LastError, Is.Null); + Assert.That((await store.Read(MigrationCategoryIds.MessageRedirects))!.State, Is.EqualTo(MigrationCategoryState.NotStarted), "a stopping host started the next category"); + Assert.That(target.RowsHandedToWrite(MigrationCategoryIds.MessageRedirects), Is.Empty); + } + } + + [Test] + public async Task A_stale_in_progress_row_the_start_has_not_reached_is_not_judged_a_stall() + { + var clock = new TimerRecordingTimeProvider(); + var store = new PollObservingCheckpointStore(clock); + var twoHoursAgo = clock.GetUtcNow().UtcDateTime - TimeSpan.FromHours(2); + await store.Upsert(new MigrationCheckpoint(MigrationCategoryIds.EndpointSettings, MigrationCategoryState.InProgress, null, 0, 0, null, null, twoHoursAgo, twoHoursAgo, null, null)); + var target = new InMemoryMigrationTarget(store) + { + DefaultBatchSize = 1, + // Each KnownEndpoints batch takes five minutes, and the next one waits until the watchdog has looked at the time. + BeforeWrite = checkpoint => + { + if (checkpoint.CategoryId == MigrationCategoryIds.KnownEndpoints) + { + clock.Advance(TimeSpan.FromMinutes(5)); + store.WaitForAPollAtTheCurrentTime().GetAwaiter().GetResult(); + } + } + }; + var source = new InMemoryMigrationSource(); + source.Seed(MigrationCategoryIds.KnownEndpoints, [.. Enumerable.Range(1, 8).Select(index => Row($"KnownEndpoints-{index}"))]); + source.Seed(MigrationCategoryIds.EndpointSettings, Row("EndpointSettings-1")); + + var results = await Task.Run(() => MigrationStartup.RunRequiredCategories( + Engine(source, target, store, clock), Categories(MigrationCategoryIds.KnownEndpoints, MigrationCategoryIds.EndpointSettings), store, clock, new CapturingLogger())) + .WaitAsync(TimeSpan.FromSeconds(30)); + + using (Assert.EnterMultipleScope()) + { + Assert.That(results.Select(checkpoint => checkpoint.State), Is.All.EqualTo(MigrationCategoryState.Complete), + "forty minutes copying KnownEndpoints were judged against a row the start had not reached yet"); + Assert.That(target.WrittenRows(MigrationCategoryIds.KnownEndpoints), Has.Count.EqualTo(8)); + Assert.That(target.WrittenRows(MigrationCategoryIds.EndpointSettings).Select(row => row.SourceId), Is.EqualTo(new[] { "EndpointSettings-1" })); + } + } + + [Test] + public void A_cancelled_copy_stays_a_cancellation() => + Assert.That(async () => await MigrationStartup.CopyOrExplainWhyItStopped(Task.FromCanceled>(new CancellationToken(canceled: true)), NewSettings()), + Throws.InstanceOf(), + "a shutdown is not a failure, so it must not come out as a refusal telling the operator to go looking at the source and the target"); + + [Test] + public void A_checkpoint_saved_by_another_instance_says_which_instance_to_stop() + { + var conflict = new MigrationCheckpointConflictException("Checkpoint KnownEndpoints was saved from version 3, but the stored row is at version 4."); + + var exception = Assert.ThrowsAsync(async () => await MigrationStartup.CopyOrExplainWhyItStopped( + Task.FromException>(conflict), + NewSettings())); + + Assert.Multiple(() => + { + Assert.That(exception!.Message, Does.Contain("a second ServiceControl pointed at the same SQLServer database"), + "the store's own message names a version, which tells nobody what is actually wrong"); + Assert.That(exception.Message, Does.Contain($"Stop the other instance, then restart with {MigrationSettings.EnabledKey} still on"), + "this is the only refusal in the subsystem where doing nothing makes it worse"); + Assert.That(exception.Message, Does.Contain(conflict.Message).And.Contain("Nothing has opened on SQLServer yet")); + Assert.That(exception.InnerException, Is.SameAs(conflict)); + }); + } + + [Test] + public async Task A_copy_that_finished_is_returned_as_it_came_back() + { + IReadOnlyList copied = [new(MigrationCategoryIds.KnownEndpoints, MigrationCategoryState.Complete, null, 3, 0, null, null, null, null, null, null)]; + + var finished = await MigrationStartup.CopyOrExplainWhyItStopped(Task.FromResult(copied), NewSettings()); + + Assert.That(finished, Is.SameAs(copied)); + } + + static MigrationRow Row(string id) => new(id, new object(), new Dictionary()); + + static MigrationCategory[] Categories(params string[] categoryIds) => [.. categoryIds.Select(categoryId => MigrationCategoryRegistry.Find(categoryId)!)]; + + static MigrationEngine Engine(InMemoryMigrationSource source, InMemoryMigrationTarget target, IMigrationCheckpointStore store, TimeProvider clock) => + new(source, target, store, clock, new MigrationEngineOptions(TimeSpan.Zero, 5, 100, []), NullLogger.Instance); + + static Settings NewSettings() => + new(transportType: "LearningTransport", persisterType: "SQLServer", errorRetentionPeriod: TimeSpan.FromDays(10)); + + // Refuses a cancelled token as the EF Core store does, so a settle made on the stalled category's token fails here too. + sealed class CancellationHonouringCheckpointStore : IMigrationCheckpointStore + { + readonly InMemoryMigrationCheckpointStore inner = new(); + + public Task> ReadAll(CancellationToken cancellationToken = default) + { + cancellationToken.ThrowIfCancellationRequested(); + return inner.ReadAll(cancellationToken); + } + + public Task Read(string categoryId, CancellationToken cancellationToken = default) + { + cancellationToken.ThrowIfCancellationRequested(); + return inner.Read(categoryId, cancellationToken); + } + + public Task Upsert(MigrationCheckpoint checkpoint, CancellationToken cancellationToken = default) + { + cancellationToken.ThrowIfCancellationRequested(); + return inner.Upsert(checkpoint, cancellationToken); + } + } +} diff --git a/src/ServiceControl.UnitTests/ScatterGather/IncompleteResultsTests.cs b/src/ServiceControl.UnitTests/ScatterGather/IncompleteResultsTests.cs index 744fa7bf19..627fee7dee 100644 --- a/src/ServiceControl.UnitTests/ScatterGather/IncompleteResultsTests.cs +++ b/src/ServiceControl.UnitTests/ScatterGather/IncompleteResultsTests.cs @@ -190,12 +190,10 @@ public async Task A_remote_timeout_does_not_cut_the_slower_remotes_short() factory.Register(settings.RemoteInstances[3], Delayed(delay, Healthy("msg-4"))); var api = new TestApi(settings, factory, Local("local-msg")); - var started = DateTime.UtcNow; var result = await api.Execute(Context(), "/api/messages"); - Assert.That(DateTime.UtcNow - started, Is.GreaterThanOrEqualTo(delay), "the composite must wait for the remotes that are still answering"); - Assert.That(result.Results.Select(m => m.MessageId), Is.EquivalentTo(["local-msg", "msg-2", "msg-3", "msg-4"])); + Assert.That(result.Results.Select(m => m.MessageId), Is.EquivalentTo(["local-msg", "msg-2", "msg-3", "msg-4"]), "the delay lives in the fake, so these three cannot arrive unless the composite waited for the remotes that are still answering"); Assert.That(result.IncompleteInstances, Is.EqualTo([new IncompleteInstance(settings.RemoteInstances[0].InstanceId, QueryFailure.TimedOut)])); Assert.That(result.QueryStats.TotalCount, Is.EqualTo(4)); } diff --git a/src/ServiceControl/HostApplicationBuilderExtensions.cs b/src/ServiceControl/HostApplicationBuilderExtensions.cs index cc43874a16..03d0d421c8 100644 --- a/src/ServiceControl/HostApplicationBuilderExtensions.cs +++ b/src/ServiceControl/HostApplicationBuilderExtensions.cs @@ -17,6 +17,7 @@ using global::ServiceControl.Operations.Metrics; using global::ServiceControl.Recoverability.Retrying.Metrics; using global::ServiceControl.Persistence; + using global::ServiceControl.Persistence.DataMigration; using global::ServiceControl.Transports; using Licensing; using Microsoft.AspNetCore.HttpLogging; @@ -103,6 +104,11 @@ public static void AddServiceControl(this IHostApplicationBuilder hostBuilder, S services.AddSingleton(provider => new Lazy(provider.GetRequiredService)); services.AddPersistence(settings); + + // The checkpoint store is optional: EF Core registers one, RavenDB does not. + services.TryAddSingleton(provider => + new CheckpointMigrationState(provider.GetService())); + services.AddMetrics(settings.PrintMetrics); hostBuilder.AddTelemetry(settings); services.AddServiceControlHealthChecks(); diff --git a/src/ServiceControl/Hosting/Commands/ErrorIngestionOnlyCommand.cs b/src/ServiceControl/Hosting/Commands/ErrorIngestionOnlyCommand.cs index 00d38bb0b3..1890ba5242 100644 --- a/src/ServiceControl/Hosting/Commands/ErrorIngestionOnlyCommand.cs +++ b/src/ServiceControl/Hosting/Commands/ErrorIngestionOnlyCommand.cs @@ -5,6 +5,7 @@ namespace ServiceControl.Hosting.Commands using System.Threading; using System.Threading.Tasks; using Microsoft.AspNetCore.Builder; + using Microsoft.Extensions.DependencyInjection; using NServiceBus; using Particular.ServiceControl; using Particular.ServiceControl.Hosting; @@ -13,8 +14,10 @@ namespace ServiceControl.Hosting.Commands using ServiceControl.ExternalIntegrations; using ServiceControl.Hosting.Https; using ServiceControl.Infrastructure.Health; + using ServiceControl.Migration; using ServiceControl.Monitoring; using ServiceControl.Persistence; + using ServiceControl.Persistence.DataMigration; using ServiceControl.Recoverability; /// @@ -25,15 +28,13 @@ namespace ServiceControl.Hosting.Commands /// class ErrorIngestionOnlyCommand : AbstractCommand { - static readonly string[] SupportedStorageNames = ["SQLServer", "PostgreSQL"]; - public override async Task Execute(HostArguments args, Settings settings, CancellationToken cancellationToken = default) { EnsureStorageCanScaleOut(settings); var app = BuildHost(settings); - await app.RunAsync(settings.RootUrl); + await app.RunAsync(); } internal static WebApplication BuildHost(Settings settings, Action customize = null) @@ -47,12 +48,27 @@ internal static WebApplication BuildHost(Settings settings, Action + new FinishedCopyBeforeAnIngestionNodeOpens( + provider.GetRequiredService(), + "this error ingestion only host")); + + if (settings.MigrationEnabled) + { + hostBuilder.Services.AddHostedService(provider => new RecordHostOpenedOnTarget(provider)); + } + customize?.Invoke(hostBuilder); var app = hostBuilder.Build(); app.MapServiceControlHealthChecks(); + // Set here rather than passed to RunAsync, so a caller that starts the host itself gets the + // configured address instead of Kestrel's default port. + app.Urls.Clear(); + app.Urls.Add(settings.RootUrl); + return app; } @@ -60,7 +76,7 @@ static void EnsureStorageCanScaleOut(Settings settings) { var manifest = PersistenceManifestLibrary.Find(settings.PersistenceType); - if (manifest == null || !SupportedStorageNames.Contains(manifest.Name, StringComparer.OrdinalIgnoreCase)) + if (manifest == null || !PersistenceFactory.SqlPersistenceNames.Contains(manifest.Name, StringComparer.OrdinalIgnoreCase)) { throw new Exception( $"--error-ingestion-only requires SQL Server or PostgreSQL storage, but this instance is configured to use '{settings.PersistenceType}'. Scaling out error ingestion is not supported for this storage type."); diff --git a/src/ServiceControl/Hosting/Commands/ImportFailedErrorsCommand.cs b/src/ServiceControl/Hosting/Commands/ImportFailedErrorsCommand.cs index af8a53840a..8e4937e504 100644 --- a/src/ServiceControl/Hosting/Commands/ImportFailedErrorsCommand.cs +++ b/src/ServiceControl/Hosting/Commands/ImportFailedErrorsCommand.cs @@ -14,6 +14,9 @@ using Recoverability; using ServiceBus.Management.Infrastructure.Settings; using ServiceControl.Infrastructure; + using ServiceControl.Migration; + using ServiceControl.Persistence; + using ServiceControl.Persistence.DataMigration; class ImportFailedErrorsCommand : AbstractCommand { @@ -54,6 +57,15 @@ internal static IHost BuildHost(Settings settings) var hostBuilder = Host.CreateApplicationBuilder(); hostBuilder.AddServiceControl(settings, endpointConfiguration, new RecoverabilityComponent()); + // Importing writes failed messages into the database, so on SQL it waits for an unfinished copy as an ingestion-only worker does. + if (PersistenceFactory.SqlPersistenceNames.Contains(PersistenceManifestLibrary.Find(settings.PersistenceType)?.Name, StringComparer.OrdinalIgnoreCase)) + { + hostBuilder.Services.AddHostedService(provider => + new FinishedCopyBeforeAnIngestionNodeOpens( + provider.GetRequiredService(), + "--import-failed-errors")); + } + return hostBuilder.Build(); } } diff --git a/src/ServiceControl/Hosting/Commands/RunCommand.cs b/src/ServiceControl/Hosting/Commands/RunCommand.cs index 33a6361511..436431f6f8 100644 --- a/src/ServiceControl/Hosting/Commands/RunCommand.cs +++ b/src/ServiceControl/Hosting/Commands/RunCommand.cs @@ -1,9 +1,12 @@ namespace ServiceControl.Hosting.Commands { + using System; using System.Threading; using System.Threading.Tasks; using Infrastructure.WebApi; using Microsoft.AspNetCore.Builder; + using Microsoft.Extensions.DependencyInjection; + using Microsoft.Extensions.Hosting; using NServiceBus; using Particular.ServiceControl; using Particular.ServiceControl.Hosting; @@ -11,11 +14,19 @@ using ServiceControl; using ServiceControl.Hosting.Auth; using ServiceControl.Hosting.Https; + using ServiceControl.Migration; using ServicePulse; class RunCommand : AbstractCommand { - public override async Task Execute(HostArguments args, Settings settings, CancellationToken cancellationToken = default) + public override Task Execute(HostArguments args, Settings settings, CancellationToken cancellationToken = default) => + Run(settings, customize: null, cancellationToken); + + /// + /// Builds and runs the full instance. When the migration is turned on, the required copy is added as the + /// first thing the host starts, so the instance opens only once the copy has finished. + /// + internal static async Task Run(Settings settings, Action customize, CancellationToken cancellationToken = default) { var endpointConfiguration = new EndpointConfiguration(settings.InstanceName); var assemblyScanner = endpointConfiguration.AssemblyScanner(); @@ -31,7 +42,17 @@ public override async Task Execute(HostArguments args, Settings settings, Cancel hostBuilder.AddServiceControl(settings, endpointConfiguration); hostBuilder.AddServiceControlApi(settings.CorsSettings); - var app = hostBuilder.Build(); + customize?.Invoke(hostBuilder); + + if (settings.MigrationEnabled) + { + hostBuilder.Services.AddHostedService(provider => new RequiredCopyBeforeTheHostOpens(provider, settings)); + hostBuilder.Services.AddHostedService(provider => new RecordHostOpenedOnTarget(provider)); + } + + // A start that refuses would otherwise leave everything the host built undisposed. + await using var app = hostBuilder.Build(); + app.UseServiceControl(settings.ForwardedHeadersSettings, settings.HttpsSettings); if (settings.EnableIntegratedServicePulse) { @@ -39,7 +60,27 @@ public override async Task Execute(HostArguments args, Settings settings, Cancel } app.UseServiceControlAuthentication(settings.OpenIdConnectSettings.Enabled); - await app.RunAsync(settings.RootUrl); + // WebApplication's RunAsync(url) takes no cancellation token, so set the url here and call IHost's RunAsync below. + app.Urls.Clear(); + app.Urls.Add(settings.RootUrl); + + // Read before RunAsync, which disposes the host before the filter below runs, so reading it there would + // throw and the catch would never fire. + var lifetime = app.Lifetime; + + // Starting the host is what marks the target as opened, through the RecordHostOpenedOnTarget hosted + // service registered above, so nothing here does it. + try + { + await app.RunAsync(cancellationToken); + } + // Stopping the service during the required copy cancels it inside host start, and that is a stop + // rather than a failure: every committed batch is durable and the next start resumes from the cursor. +#pragma warning disable PS0020 // The host cancels on its own lifetime token, not the caller's, so that is the one to filter on + catch (OperationCanceledException) when (settings.MigrationEnabled && lifetime.ApplicationStopping.IsCancellationRequested) +#pragma warning restore PS0020 + { + } } } } diff --git a/src/ServiceControl/Hosting/Help.txt b/src/ServiceControl/Hosting/Help.txt index 6fba5341f9..d616180539 100644 --- a/src/ServiceControl/Hosting/Help.txt +++ b/src/ServiceControl/Hosting/Help.txt @@ -23,20 +23,6 @@ This mode runs no setup, so combining it with --setup or --setup-and-run is refu Message bodies must be stored somewhere every host can read, so this mode should not be combined with file system body storage unless the path is a shared mount. -MIGRATION SOURCE REPORT - - ServiceControl.exe --migration-source-report - -Reports what a migration would read from the RavenDB source: the facts the source reports about -itself, with the setting each came from, and a row count for everything it holds. The source reads the instance's own -RavenDB settings, so keep them in place when switching PersistenceType. - -For a RavenDB source: an EXTERNAL server can be reported on while ServiceControl is running. An EMBEDDED source -cannot: ServiceControl starts its own RavenDB process against that data directory, and a second one -cannot attach to it. Stop the ServiceControl service first, run the report, and start it again. An -instance that does not ship the RavenDB server, such as the container image, cannot report on an -embedded source at all. - SERVICE INSTALL AND UNINSTALL AND CONFIGURATION OPTIONS As of Service Control 1.7 the command line uninstall and install switches have been removed. diff --git a/src/ServiceControl/Infrastructure/Settings/Settings.cs b/src/ServiceControl/Infrastructure/Settings/Settings.cs index af52f4bf88..35f2c65d75 100644 --- a/src/ServiceControl/Infrastructure/Settings/Settings.cs +++ b/src/ServiceControl/Infrastructure/Settings/Settings.cs @@ -16,6 +16,7 @@ using ServiceControl.Infrastructure.Settings; using ServiceControl.Infrastructure.WebApi; using ServiceControl.Persistence; + using ServiceControl.Persistence.DataMigration; using ServiceControl.Transports; using ServicePulse; using JsonSerializer = System.Text.Json.JsonSerializer; @@ -185,6 +186,7 @@ public string InstanceId public string TransportType { get; set; } public string PersistenceType { get; private set; } + public bool MigrationEnabled => SettingsReader.Read(SettingsRootNamespace, MigrationSettings.EnabledKey, MigrationSettings.DefaultEnabled); public string ErrorLogQueue { get; set; } public string ErrorQueue { get; set; } diff --git a/src/ServiceControl/Migration/Checks/MigrationIsReleasedCheck.cs b/src/ServiceControl/Migration/Checks/MigrationIsReleasedCheck.cs new file mode 100644 index 0000000000..68ab286716 --- /dev/null +++ b/src/ServiceControl/Migration/Checks/MigrationIsReleasedCheck.cs @@ -0,0 +1,32 @@ +namespace ServiceControl.Migration.Checks; + +using System; +using System.Threading; +using System.Threading.Tasks; +using ServiceControl.Persistence.DataMigration; + +/// +/// Registered by a test host that needs a copy to run on a build that does not yet carry the whole migration. +/// +class AllowUnreleasedMigration; + +/// +/// Refuses every copy until the whole migration has shipped. A build that can copy the required categories but +/// not yet the background copy, the end-of-migration guard or verification would still commit an instance to the +/// target with no way to finish or check the move. The last phase of the migration deletes this check. +/// +class MigrationIsReleasedCheck(AllowUnreleasedMigration allowUnreleasedMigration = null) : IMigrationStartupCheck +{ + public string Name => "this build carries the whole migration"; + + public Task Run(CancellationToken cancellationToken = default) + { + if (allowUnreleasedMigration is null) + { + throw new Exception( + $"This build of ServiceControl does not yet carry the whole migration from RavenDB, so it will not copy anything. Set {MigrationSettings.EnabledKey} back to false and start ServiceControl again."); + } + + return Task.CompletedTask; + } +} diff --git a/src/ServiceControl/Migration/Checks/MigrationPairIsSupportedCheck.cs b/src/ServiceControl/Migration/Checks/MigrationPairIsSupportedCheck.cs new file mode 100644 index 0000000000..e3ffed392a --- /dev/null +++ b/src/ServiceControl/Migration/Checks/MigrationPairIsSupportedCheck.cs @@ -0,0 +1,31 @@ +namespace ServiceControl.Migration.Checks; + +using System; +using System.Linq; +using System.Threading; +using System.Threading.Tasks; +using ServiceBus.Management.Infrastructure.Settings; +using ServiceControl.Persistence; +using ServiceControl.Persistence.DataMigration; + +/// +/// Refuses any target but SQL Server or PostgreSQL, since the source is always RavenDB. The seams underneath are +/// general enough to copy between any two persisters, and this is the check that says which pair is actually supported. +/// +class MigrationPairIsSupportedCheck(Settings settings) : IMigrationStartupCheck +{ + public string Name => "the source and target are the supported pair"; + + public Task Run(CancellationToken cancellationToken = default) + { + var targetName = PersistenceManifestLibrary.Find(settings.PersistenceType)?.Name; + + if (!PersistenceFactory.SqlPersistenceNames.Contains(targetName, StringComparer.OrdinalIgnoreCase)) + { + throw new Exception( + $"Migrating from '{PersistenceFactory.MigrationSourcePersistenceType}' to '{settings.PersistenceType}' is not supported. The only supported migration is from RavenDB to SQL Server or PostgreSQL, so set {Settings.SettingsRootNamespace}/PersistenceType to {string.Join(" or ", PersistenceFactory.SqlPersistenceNames)} before setting {MigrationSettings.EnabledKey}."); + } + + return Task.CompletedTask; + } +} diff --git a/src/ServiceControl/Migration/Checks/OptionalCategoryWindowsAreValidCheck.cs b/src/ServiceControl/Migration/Checks/OptionalCategoryWindowsAreValidCheck.cs new file mode 100644 index 0000000000..d8f2ef87bc --- /dev/null +++ b/src/ServiceControl/Migration/Checks/OptionalCategoryWindowsAreValidCheck.cs @@ -0,0 +1,27 @@ +namespace ServiceControl.Migration.Checks; + +using System; +using System.Threading; +using System.Threading.Tasks; +using ServiceBus.Management.Infrastructure.Settings; +using ServiceControl.Persistence.DataMigration; + +/// +/// Reads the engine options out of the settings, which refuses an optional category window that is not a time +/// span before anything is copied rather than part way through the copy. +/// +class OptionalCategoryWindowsAreValidCheck(TimeSpan eventRetentionPeriod, TimeSpan errorRetentionPeriod) : IMigrationStartupCheck +{ + public string Name => "the optional category windows are valid"; + + /// + /// The options parsed from the settings, which is null until has returned. + /// + public MigrationEngineOptions Options { get; private set; } + + public Task Run(CancellationToken cancellationToken = default) + { + Options = MigrationEngineOptions.FromSettings(Settings.SettingsRootNamespace, eventRetentionPeriod, errorRetentionPeriod); + return Task.CompletedTask; + } +} diff --git a/src/ServiceControl/Migration/Checks/RetryHistoryDepthIsSafeCheck.cs b/src/ServiceControl/Migration/Checks/RetryHistoryDepthIsSafeCheck.cs new file mode 100644 index 0000000000..00f5903bd4 --- /dev/null +++ b/src/ServiceControl/Migration/Checks/RetryHistoryDepthIsSafeCheck.cs @@ -0,0 +1,28 @@ +namespace ServiceControl.Migration.Checks; + +using System; +using System.Threading; +using System.Threading.Tasks; +using ServiceControl.Persistence.DataMigration; + +/// +/// Refuses a copy that the instance's own retry history setting would throw away. At a depth of zero the SQL +/// persisters empty the history table whenever a retry completes, and the migrated rows would go with it. +/// +/// The instance's ServiceControl/RetryHistoryDepth. +class RetryHistoryDepthIsSafeCheck(int retryHistoryDepth) : IMigrationStartupCheck +{ + public string Name => "the retry history depth will not empty a migrated table"; + + public Task Run(CancellationToken cancellationToken = default) + { + // The depth at which RetryHistoryDataStore.TrimHistory deletes every row rather than trimming. + if (retryHistoryDepth <= 0) + { + throw new Exception( + $"ServiceControl/RetryHistoryDepth is {retryHistoryDepth}, and at that depth the SQL persisters delete every row of HistoricRetryOperations when a retry completes. The RetryOperations category copies that table, so the first completed retry after cutover would silently discard the migrated retry history. Set ServiceControl/RetryHistoryDepth to a positive number, or clear it to use the default of 10, before setting {MigrationSettings.EnabledKey}."); + } + + return Task.CompletedTask; + } +} diff --git a/src/ServiceControl/Migration/FinishedCopyBeforeAnIngestionNodeOpens.cs b/src/ServiceControl/Migration/FinishedCopyBeforeAnIngestionNodeOpens.cs new file mode 100644 index 0000000000..6433658d4b --- /dev/null +++ b/src/ServiceControl/Migration/FinishedCopyBeforeAnIngestionNodeOpens.cs @@ -0,0 +1,58 @@ +namespace ServiceControl.Migration; + +using System; +using System.Linq; +using System.Threading; +using System.Threading.Tasks; +using Microsoft.Extensions.Hosting; +using ServiceControl.Persistence.DataMigration; + +/// +/// Refuses to start a host that writes failed messages into the database, an error ingestion only host or +/// --import-failed-errors, until every required category is Done or Abandoned. An optional category, which copies +/// in the background once the main host has opened, never holds it back, but a category under an id this build +/// does not know does. A refusal throws, which fails the start and stops the host. Start such a host only after the +/// main host has opened, because a database with no checkpoint row yet reads as one no migration has touched. +/// +/// Where the checkpoint rows are read. A database no migration has touched holds none, so the host starts as before. +/// Names the host in the refusal, such as "this error ingestion only host". +sealed class FinishedCopyBeforeAnIngestionNodeOpens( + IMigrationCheckpointStore checkpointStore, + string hostDescription) : IHostedLifecycleService +{ + // It runs with the migration off too, because a node someone forgot to flag would ingest into a part-copied + // database and turn abandoning it from a clean rollback into permanent loss. A database that never migrated has + // nothing unfinished, so its nodes start as before. + public async Task StartingAsync(CancellationToken cancellationToken = default) + { + var unfinished = (await checkpointStore.ReadAll(cancellationToken)) + .Where(checkpoint => MigrationCategoryRegistry.Find(checkpoint.CategoryId)?.Kind != MigrationCategoryKind.Optional) + .Where(checkpoint => !checkpoint.State.IsFinished()) + .ToArray(); + + if (unfinished.Length == 0) + { + return; + } + + var detail = string.Join(". ", unfinished.Select(checkpoint => checkpoint.State.IsFailed() + ? MigrationStartup.FailedDetail(checkpoint) + : $"{checkpoint.CategoryId} is Copying ({checkpoint.State})")); + + var advice = unfinished.Any(checkpoint => !checkpoint.State.IsFailed()) + ? "Let the instance running the copy finish it and start this host again." + : "Start this host again once they are settled."; + + throw new Exception($"A copy into this database has not finished, so {hostDescription} will not start and nothing has been lost. {detail}. {advice}"); + } + + public Task StartAsync(CancellationToken cancellationToken = default) => Task.CompletedTask; + + public Task StartedAsync(CancellationToken cancellationToken = default) => Task.CompletedTask; + + public Task StoppingAsync(CancellationToken cancellationToken = default) => Task.CompletedTask; + + public Task StopAsync(CancellationToken cancellationToken = default) => Task.CompletedTask; + + public Task StoppedAsync(CancellationToken cancellationToken = default) => Task.CompletedTask; +} diff --git a/src/ServiceControl/Migration/MigrationStartup.cs b/src/ServiceControl/Migration/MigrationStartup.cs new file mode 100644 index 0000000000..dc17bab1d1 --- /dev/null +++ b/src/ServiceControl/Migration/MigrationStartup.cs @@ -0,0 +1,607 @@ +namespace ServiceControl.Migration; + +using System; +using System.Collections.Generic; +using System.Linq; +using System.Threading; +using System.Threading.Tasks; +using Microsoft.Extensions.DependencyInjection; +using Microsoft.Extensions.Logging; +using ServiceBus.Management.Infrastructure.Settings; +using ServiceControl.Migration.Checks; +using ServiceControl.Persistence; +using ServiceControl.Persistence.DataMigration; + +/// +/// The required copy, from the checks that decide whether it can run at all to the refusal it throws when a +/// category did not finish. Everything here happens with ServiceControl closed, which is the only window in +/// which the copy can be thrown away at no cost. +/// +static class MigrationStartup +{ + // A category needs both a reader and a writer, so a half-implemented one is never attempted. + internal static IReadOnlyCollection CopyableCategoryIds(IReadOnlyCollection sourceSupports, IReadOnlyCollection targetSupports) => + [.. sourceSupports.Intersect(targetSupports, StringComparer.Ordinal)]; + + /// + /// Runs the startup checks, copies every required category this build can copy that is not already settled, + /// and then seeds the the host reads. Once every required category this build + /// copies is settled and no row under an unknown id is unfinished, a source that will not open is logged and + /// recorded on the optional categories' rows instead of refused, nothing is copied, and the host still opens. Throws when a check refuses or a category does not + /// finish, with a message telling the operator what to do and how to go back; the caller must let that stop + /// the host. Call it after the host is built and before it starts, because the copy has to finish before + /// anything else opens on the target. + /// + /// The built host's services, which is where the target, the checkpoint store and the migration state come from. + /// The instance settings, read for the source and target persistence types. + /// Cancelled when the host is shutting down, which ends the copy without a refusal message. + public static async Task RunRequiredCopy(IServiceProvider services, Settings settings, CancellationToken cancellationToken = default) + { + await MigrationStartupCheckRunner.Run( + [ + new MigrationIsReleasedCheck(services.GetService()), + new MigrationPairIsSupportedCheck(settings) + ], cancellationToken); + + var loggerFactory = services.GetRequiredService(); + var logger = loggerFactory.CreateLogger(typeof(MigrationStartup)); + var target = services.GetRequiredService(); + var checkpointStore = services.GetRequiredService(); + var timeProvider = services.GetRequiredService(); + + await using var source = PersistenceFactory.CreateMigrationSource(settings); + + var copyable = CopyableCategoryIds(source.SupportedCategoryIds, target.SupportedCategoryIds); + + var options = await RunChecksAndOpenTarget(services, settings, cancellationToken); + + var engine = new MigrationEngine( + source, + target, + checkpointStore, + timeProvider, + options, + loggerFactory.CreateLogger()); + + // Optional categories copy beside the running services, so only the required ones hold those services back. + var holdingBack = engine.SelectCategories(MigrationCategoryKind.Required) + .Select(category => category.Id) + .ToArray(); + + var toCopy = engine.SelectCategories(MigrationCategoryKind.Required) + .Where(category => copyable.Contains(category.Id)) + .ToArray(); + + var deferred = engine.SelectCategories(MigrationCategoryKind.Required) + .Where(category => !copyable.Contains(category.Id)) + .Select(category => category.Id) + .ToArray(); + + var checkpoints = await checkpointStore.ReadAll(cancellationToken); + var unfinishedOutsideTheCopy = UnfinishedRowsOutsideTheCopy(checkpoints, toCopy); + var requiredCopySettled = RequiredCopyIsSettled(checkpoints, toCopy); + + try + { + await MigrationStartupCheckRunner.Run([SourceOpens(source)], cancellationToken); + } + catch (OperationCanceledException) when (cancellationToken.IsCancellationRequested) + { + throw; + } + // Safe to open without the source, because every required category is settled and only the background copy of the optional ones reads it. + catch (Exception exception) when (requiredCopySettled) + { + // The runner's wrapper says ServiceControl will not start, which is false here and would land on the rows. + var cause = exception.InnerException ?? exception; + + logger.LogError(cause, + "The {SourcePersistenceType} migration source could not be opened, so ServiceControl opens without it: every required category this build copies is Done or Abandoned and no row under an id it does not know is unfinished, so only the background copy of the optional categories reads the source. The optional categories still copying stay as they are until a start that can open the source, and their checkpoints carry this error.", + PersistenceFactory.MigrationSourcePersistenceType); + + await RecordSourceOutage(checkpointStore, engine.SelectCategories(MigrationCategoryKind.Optional), cause, cancellationToken); + await SeedMigrationState(services, holdingBack, cancellationToken); + return; + } + catch (Exception exception) + { + throw new Exception($"{exception.Message} {NoWayOutWithoutTheSource(settings)}", exception.InnerException); + } + + await MigrationStartupCheckRunner.Run(source.ContributedChecks(), cancellationToken); + + logger.LogInformation( + "Migration mode: copying {CopyCount} required categories before ServiceControl opens ({DeferredCount} not yet implemented: {Deferred})", + toCopy.Length, deferred.Length, string.Join(", ", deferred)); + + // Captured before the copy so the report can tell a category this run finished from one an earlier run did. + var runStartedAt = timeProvider.GetUtcNow().UtcDateTime; + + var finished = await CopyOrExplainWhyItStopped( + RunRequiredCategories(engine, toCopy, checkpointStore, timeProvider, logger, cancellationToken), + settings); + + ReportWhatTheCopyLeftBehind(finished, logger, runStartedAt); + + RefuseIfAnyCategoryDidNotComplete(toCopy, finished, unfinishedOutsideTheCopy, settings); + + await SeedMigrationState(services, holdingBack, cancellationToken); + } + + static async Task SeedMigrationState(IServiceProvider services, IReadOnlyCollection holdingBack, CancellationToken cancellationToken) + { + if (services.GetRequiredService() is CheckpointMigrationState state) + { + await state.Seed(holdingBack, cancellationToken); + } + } + + /// + /// Runs every startup check in the order they have to run, and opens the target and the source as two of + /// them. The order is what the operator sees: a check that costs nothing comes before one that connects to a + /// database, and the source's own checks run last because they need it open. + /// + /// The instance settings, whose retention periods are the optional category windows when none is set, and whose retry history depth the copy must not be thrown away by. + /// The options read from the settings, which the window check parsed on its way past. + /// A check refused. The message names the check and says what to do, and this start has copied nothing. + public static async Task RunChecksAndOpen(IServiceProvider services, Settings settings, IMigrationSource source, CancellationToken cancellationToken = default) + { + var options = await RunChecksAndOpenTarget(services, settings, cancellationToken); + + await MigrationStartupCheckRunner.Run([SourceOpens(source)], cancellationToken); + + await MigrationStartupCheckRunner.Run(source.ContributedChecks(), cancellationToken); + + return options; + } + + /// + /// Runs every startup check that does not need the source, in the order they have to run, ending with + /// opening the target. A check that costs nothing comes before one that connects to a database. The target + /// is open when this returns, and the source has not been touched. + /// + /// The built host's services, which is where the target and its readiness checks come from. + /// The instance settings, whose retention periods are the optional category windows when none is set, and whose retry history depth the copy must not be thrown away by. + /// Cancelled when the host is shutting down, which stops the checks without a refusal message. + /// The options read from the settings, which the window check parsed on its way past. + /// A check refused or the target did not open. The message names the check and says what to do, and this start has copied nothing. + /// The host is shutting down. + public static async Task RunChecksAndOpenTarget(IServiceProvider services, Settings settings, CancellationToken cancellationToken = default) + { + var target = services.GetRequiredService(); + var readiness = services.GetRequiredService(); + var categories = new OptionalCategoryWindowsAreValidCheck(settings.EventsRetentionPeriod, settings.ErrorRetentionPeriod); + + await MigrationStartupCheckRunner.Run( + [ + categories, + new RetryHistoryDepthIsSafeCheck(settings.RetryHistoryDepth), + .. readiness.ContributedChecks(), + new Step("the migration target opens", target.Open) + ], cancellationToken); + + return categories.Options; + } + + /// + /// Records on each optional category's checkpoint that the source could not be opened, so status and verify + /// read the outage from the rows. The required copy calls it on a start that + /// opens without the source, and the background copier calls it when its own open fails. It changes no + /// state, cursor or count: a row still copying gains the error as its , + /// a category with no row gets a new not-started row carrying the error, and a row that is Done, Failed or + /// Abandoned is left alone. The next start that resumes a row clears the error. + /// + /// Where the rows are read and saved. + /// The optional categories this instance selected. + /// Why the source could not be opened. Its type and message go into the error. + /// Cancelled when the host is shutting down. + /// Another writer saved one of these rows since it was read. + internal static async Task RecordSourceOutage(IMigrationCheckpointStore checkpointStore, IReadOnlyList optionalCategories, Exception exception, CancellationToken cancellationToken = default) + { + var error = $"The {PersistenceFactory.MigrationSourcePersistenceType} migration source could not be opened: {exception.GetType().Name}: {exception.Message.TrimEnd('.', ' ')}. This category stays as it is, and the next start tries the source again."; + + foreach (var category in optionalCategories) + { + var checkpoint = await checkpointStore.Read(category.Id, cancellationToken); + + if (checkpoint is null) + { + await checkpointStore.Upsert(new MigrationCheckpoint(category.Id, MigrationCategoryState.NotStarted, null, 0, 0, null, null, null, null, null, error), cancellationToken); + } + else if (!checkpoint.State.IsFinished() && !checkpoint.State.IsFailed()) + { + await checkpointStore.Upsert(checkpoint with { LastError = error }, cancellationToken); + } + } + } + + /// + /// Copies the required categories one after another under the stall watchdog and returns where each one ended, + /// in the same order. A category that commits nothing for is + /// stopped and settled Halted, and the categories after it still get their go. Before copying anything it saves + /// a not-started checkpoint for every category that has none, as does. + /// + /// Copies each category. + /// The categories to copy, in the order to copy them. + /// Where each category's row is seeded, watched and, after a stall, settled. + /// The clock the watchdog measures a stall on. + /// Receives the watchdog's progress lines and the stall. + /// The host's token. Cancelling it stops the copy and leaves the running category as its last committed batch left it, so the next start resumes it. + /// The checkpoint each category ended on. + /// Another writer saved one of these checkpoints, which means a second instance is copying into the same database. + /// The host is shutting down. + internal static async Task> RunRequiredCategories( + MigrationEngine engine, + IReadOnlyList categories, + IMigrationCheckpointStore checkpointStore, + TimeProvider timeProvider, + ILogger logger, + CancellationToken cancellationToken = default) + { + // The gates that keep a host off an unfinished copy read only the rows that exist. + foreach (var category in categories) + { + if (await checkpointStore.Read(category.Id, cancellationToken) is null) + { + await checkpointStore.Upsert(new MigrationCheckpoint(category.Id, MigrationCategoryState.NotStarted, null, 0, 0, null, null, null, null, null, null), cancellationToken); + } + } + + await using var progress = new ClosedWindowProgress(checkpointStore, timeProvider, logger, [.. categories.Select(category => category.Id)], cancellationToken); + + var results = new List(categories.Count); + + foreach (var category in categories) + { + using var categoryCancellation = CancellationTokenSource.CreateLinkedTokenSource(cancellationToken); + progress.Watch(category.Id, timeProvider.GetUtcNow().UtcDateTime, categoryCancellation); + + try + { + results.Add(await engine.RunCategoryAsync(category, categoryCancellation.Token)); + } + // Only the watchdog cancels this source without the host's token, so the filter tells a stall from a shutdown. +#pragma warning disable PS0020 + catch (OperationCanceledException) when (categoryCancellation.IsCancellationRequested && !cancellationToken.IsCancellationRequested) +#pragma warning restore PS0020 + { + // The host's token, because the stall has already cancelled the category's and a store call on that would fail. + var stalled = await checkpointStore.Read(category.Id, cancellationToken); + + results.Add(await checkpointStore.Upsert(stalled with + { + State = MigrationCategoryState.Halted, + SettledAt = timeProvider.GetUtcNow().UtcDateTime, + LastError = StallExplanation(category.Id) + }, cancellationToken)); + } + } + + return results; + } + + /// + /// Waits for the copy and turns a second instance writing checkpoints to the same database into a refusal the + /// operator can act on. A shutdown and every other failure come out as they are. + /// + /// The copy, already running. + /// Read for the persistence type the refusal names. + /// What the copy returned. + /// A checkpoint saved by another instance. The message says what to do and how to go back. + /// The host is shutting down. +#pragma warning disable PS0018 // The copy it waits on already runs under the host's token, so a token here would have nothing to cancel. + internal static async Task> CopyOrExplainWhyItStopped( + Task> copy, + Settings settings) +#pragma warning restore PS0018 + { + try + { + return await copy; + } + // The engine never turns a conflict into a halt, so without this the copy ends on the checkpoint store's + // own message and none of the advice every other refusal carries. + catch (MigrationCheckpointConflictException exception) + { + throw new Exception( + $"The required copy stopped because another writer saved a migration checkpoint for this instance, which is a second ServiceControl pointed at the same {settings.PersistenceType} database. ServiceControl will not start. {exception.Message} Stop the other instance, then restart with {MigrationSettings.EnabledKey} still on; the copy resumes from its last committed batch. " + + RollbackAdvice(settings), exception); + } + } + + /// + /// Throws unless every attempted category finished. A category that reported no checkpoint at all counts as + /// outstanding too, because nothing says how much of it was copied. + /// + /// A category did not finish. The message names each one with its state and counts, names the commands that move each Failed one on, and says how to go back. + internal static void RefuseIfAnyCategoryDidNotComplete( + IReadOnlyList attempted, + IReadOnlyList finished, + Settings settings) + { + var outstanding = finished + .Where(checkpoint => !checkpoint.State.IsFinished()) + .Select(checkpoint => checkpoint.State.IsFailed() ? FailedDetail(checkpoint) : CopyingDetail(checkpoint)) + .ToList(); + + var reported = finished.Select(checkpoint => checkpoint.CategoryId).ToHashSet(StringComparer.Ordinal); + + outstanding.AddRange(attempted + .Where(category => !reported.Contains(category.Id)) + .Select(category => $"{category.Id} reported no checkpoint at all, so whether it copied anything is unknown")); + + if (outstanding.Count == 0) + { + return; + } + + throw new Exception( + $"The required copy did not finish, so ServiceControl will not start and nothing has been lost. {string.Join(". ", outstanding)}. " + + RollbackAdvice(settings)); + } + + /// + /// Says whether the host may open without the source: every category this start copies is finished, and no row + /// outside the copy is unfinished. + /// + /// Every checkpoint row in the target. + /// The categories this start copies. + /// True when nothing required is still outstanding. + internal static bool RequiredCopyIsSettled(IReadOnlyList checkpoints, IReadOnlyList toCopy) => + UnfinishedRowsOutsideTheCopy(checkpoints, toCopy).Count == 0 + && toCopy.All(category => checkpoints.Any(checkpoint => checkpoint.CategoryId == category.Id && checkpoint.State.IsFinished())); + + /// + /// Throws unless every attempted category finished and no row outside the copy is unfinished, naming each + /// outstanding row in the same words. + /// + /// The categories this start copied. + /// Where each attempted category ended. + /// The rows found. + /// Read for the persistence type the refusal names. + /// Something is outstanding. The message names it and says how to go back. + internal static void RefuseIfAnyCategoryDidNotComplete( + IReadOnlyList attempted, + IReadOnlyList finished, + IReadOnlyList unfinishedOutsideTheCopy, + Settings settings) => + RefuseIfAnyCategoryDidNotComplete(attempted, [.. finished, .. unfinishedOutsideTheCopy], settings); + + /// + /// Finds the unfinished rows the copy will not run: a row under an id this build does not know, which counts as + /// required because a newer build may have written it for a category that must finish, and a required row for a + /// category this build cannot copy yet. A finished row, and a row known to be optional, hold nothing. + /// + /// Every checkpoint row in the target. + /// The categories this start copies, whose rows the copy judges itself. + /// The rows that keep the host closed, in the order they were read. + internal static IReadOnlyList UnfinishedRowsOutsideTheCopy(IReadOnlyList checkpoints, IReadOnlyList toCopy) => + [.. checkpoints + .Where(checkpoint => MigrationCategoryRegistry.Find(checkpoint.CategoryId)?.Kind != MigrationCategoryKind.Optional) + .Where(checkpoint => toCopy.All(category => category.Id != checkpoint.CategoryId)) + .Where(checkpoint => !checkpoint.State.IsFinished())]; + + internal static string FailedDetail(MigrationCheckpoint checkpoint) + { + var reasons = checkpoint.SkipReasons is { Count: > 0 } counts + ? $" ({string.Join("; ", counts.Select(reason => $"{reason.Key} {reason.Value}{(reason.Key.IsPermanent() ? ", no retry can fix" : "")}"))})" + : ""; + + return $"{checkpoint.CategoryId} is Failed ({checkpoint.State}) after copying {checkpoint.CopiedCount} and skipping {checkpoint.SkippedCount}{reasons}{LastErrorClause(checkpoint)}; " + + $"run --migration-retry {checkpoint.CategoryId} once the cause is fixed, which copies the category again from the start, or --migration-abandon {checkpoint.CategoryId} to keep what was copied and give up the rest, both with ServiceControl stopped"; + } + + static string CopyingDetail(MigrationCheckpoint checkpoint) => + $"{checkpoint.CategoryId} is {checkpoint.State} after copying {checkpoint.CopiedCount} and skipping {checkpoint.SkippedCount}{LastErrorClause(checkpoint)}"; + + // Trimmed because the refusal punctuates each category's sentence itself. + static string LastErrorClause(MigrationCheckpoint checkpoint) => + checkpoint.LastError is null ? "" : $": {checkpoint.LastError.TrimEnd('.', ' ')}"; + + /// + /// Logs one line per category saying what it copied and what it left behind. Skipped rows are warned about + /// one at a time while the copy runs, and nothing else states the total or says what becomes of them. + /// + /// When this start began, which is what tells a category this run finished from one an earlier run did. + internal static void ReportWhatTheCopyLeftBehind(IReadOnlyList finished, ILogger logger, DateTime runStartedAt) + { + foreach (var checkpoint in finished) + { + // It did not run, and the refusal that follows names it with the category it waits for. + if (checkpoint.State == MigrationCategoryState.Blocked) + { + continue; + } + + // A category already finished when this run began is returned without being run, so its counts are + // an earlier run's. A Failed one is returned untouched too, but it is not finished, so it falls through. + if (checkpoint.State.IsFinished() && checkpoint.SettledAt is { } settledAt && settledAt < runStartedAt) + { + logger.LogInformation( + "{CategoryId}: already finished before this start, by a run that copied {Copied} and skipped {Skipped}. This start copied nothing.", + checkpoint.CategoryId, checkpoint.CopiedCount, checkpoint.SkippedCount); + continue; + } + + // An exception, a stall or unbalanced counts fail a category without a skip, and that is not a clean copy. + if (checkpoint.State.IsFailed() && checkpoint.SkippedCount == 0) + { + logger.LogWarning("{CategoryId} is Failed ({State}) after {Copied} copied and {AlreadyPresent} already present. The refusal that follows says why and what to run.", + checkpoint.CategoryId, checkpoint.State, checkpoint.CopiedCount, checkpoint.AlreadyPresentCount); + continue; + } + + if (checkpoint.SkippedCount == 0) + { + logger.LogInformation("{CategoryId}: {Copied} copied, {AlreadyPresent} already present, nothing skipped", + checkpoint.CategoryId, checkpoint.CopiedCount, checkpoint.AlreadyPresentCount); + continue; + } + + var reasons = checkpoint.SkipReasons is { Count: > 0 } counts + ? string.Join(", ", counts.OrderByDescending(reason => reason.Value).Select(reason => $"{reason.Key} {reason.Value}")) + : "no reason recorded"; + + // Only a Failed category can be retried, so a Done one is never offered the command. + var whatBecomesOfThem = checkpoint.State.IsFailed() + ? WhatARetryCanDo(checkpoint) + : checkpoint.State == MigrationCategoryState.Complete + ? "They were left out as harmless, because ServiceControl would have removed them anyway, and they stay only in the source database." + : "They stay only in the source database."; + + logger.LogWarning( + "{CategoryId}: {Copied} copied, {AlreadyPresent} already present, {Skipped} skipped ({Reasons}). {WhatBecomesOfThem}", + checkpoint.CategoryId, checkpoint.CopiedCount, checkpoint.AlreadyPresentCount, checkpoint.SkippedCount, reasons, whatBecomesOfThem); + } + } + + // A retry re-reads the whole category, so it is offered only when some of the faults could come across on it. + static string WhatARetryCanDo(MigrationCheckpoint checkpoint) + { + var faults = checkpoint.SkipReasons?.Keys.Where(reason => !reason.IsBenign()).ToArray() ?? []; + var permanent = faults.Where(reason => reason.IsPermanent()).ToArray(); + + if (faults.Length > 0 && permanent.Length == faults.Length) + { + return $"They stay only in the source database, and no retry can fix them, so --migration-abandon {checkpoint.CategoryId} keeps what was copied and gives up the rest."; + } + + var cannotFix = permanent.Length == 0 ? "" : $" No retry can fix the {string.Join(" or ", permanent)} ones."; + + return $"They stay only in the source database until --migration-retry {checkpoint.CategoryId} re-reads them once the cause is fixed.{cannotFix}"; + } + + internal static string StallExplanation(string stalledCategoryId) => + $"Halted: {stalledCategoryId} committed nothing for {ClosedWindowProgress.StallLimit.TotalMinutes:0.#} minutes, so it was stopped. The limit is not configurable: check that the source and the target are both responding rather than looking for a setting to change."; + + static string RollbackAdvice(Settings settings) => + $"Nothing has opened on {settings.PersistenceType} yet, so setting {MigrationSettings.EnabledKey}=false and pointing PersistenceType back at {PersistenceFactory.MigrationSourcePersistenceType} discards the partial copy and returns the instance to {PersistenceFactory.MigrationSourcePersistenceType} with no loss."; + + static string NoWayOutWithoutTheSource(Settings settings) => + $"If the {PersistenceFactory.MigrationSourcePersistenceType} source is already gone, a required category that is Failed or has started copying can be given up with --migration-abandon , with ServiceControl stopped. " + + $"A required category that never started cannot be abandoned: while one is outstanding ServiceControl has never opened on {settings.PersistenceType}, so pointing PersistenceType back at {PersistenceFactory.MigrationSourcePersistenceType}, or starting over against an empty {settings.PersistenceType} database, loses nothing it has served."; + + static Step SourceOpens(IMigrationSource source) => new("the migration source opens", source.Open); + + /// + /// Makes opening the target or the source look like a startup check, so a failure to connect is reported in + /// the same words as a check that refused, and in its place in the order. + /// + sealed class Step(string name, Func run) : IMigrationStartupCheck + { + public string Name => name; + + public Task Run(CancellationToken cancellationToken = default) => run(cancellationToken); + } + + /// + /// Logs how far each category has got while the copy runs, and stops the running category when it commits + /// nothing for . Without it a copy that is waiting on a database nobody is watching + /// holds the instance closed for as long as the operator leaves it. + /// + internal sealed class ClosedWindowProgress : IAsyncDisposable + { + // Not configurable, and not a total timeout: a deadline would kill a copy that is working. + internal static readonly TimeSpan PollInterval = TimeSpan.FromSeconds(30); + internal static readonly TimeSpan StallLimit = TimeSpan.FromMinutes(30); + + readonly CancellationTokenSource cancellation; + readonly HashSet attempted; + readonly Task polling; + volatile RunningCategory running; + + public ClosedWindowProgress(IMigrationCheckpointStore checkpointStore, TimeProvider timeProvider, ILogger logger, IReadOnlyCollection attemptedCategoryIds, CancellationToken cancellationToken = default) + { + cancellation = CancellationTokenSource.CreateLinkedTokenSource(cancellationToken); + attempted = attemptedCategoryIds.ToHashSet(StringComparer.Ordinal); + polling = Poll(checkpointStore, timeProvider, logger, cancellation.Token); + } + + /// + /// Points the watchdog at the category that is starting now, in place of the one before it. Only this + /// category is judged from here on, from the later of its last committed batch and . + /// + /// The category starting now. + /// When this start began running it. A row an earlier start left carries an older stamp, which would read as a stall at once. + /// Cancelled when the category stalls. The category must run under its token, and nothing else should. + public void Watch(string categoryId, DateTime runStartedAt, CancellationTokenSource stop) => + running = new RunningCategory(categoryId, runStartedAt, stop); + + async Task Poll(IMigrationCheckpointStore checkpointStore, TimeProvider timeProvider, ILogger logger, CancellationToken cancellationToken) + { + using var timer = new PeriodicTimer(PollInterval, timeProvider); + + try + { + while (await timer.WaitForNextTickAsync(cancellationToken)) + { + try + { + // Only this run's categories, because a row an earlier run left in progress is not this run's to report. + var inProgress = (await checkpointStore.ReadAll(cancellationToken)) + .Where(checkpoint => checkpoint.State == MigrationCategoryState.InProgress && attempted.Contains(checkpoint.CategoryId)) + .ToArray(); + + foreach (var checkpoint in inProgress) + { + logger.LogInformation( + "{CategoryId}: {Copied} of {Total} copied, {Skipped} skipped, cursor {Cursor}", + checkpoint.CategoryId, + checkpoint.CopiedCount, + checkpoint.SourceTotal is { } total ? total.ToString() : "an unknown number of", + checkpoint.SkippedCount, + checkpoint.Cursor ?? "the start"); + } + + var watched = running; + + // A resumed row carries the previous run's stamp, so the window starts at whichever is + // later: that stamp, or the moment this category's run began. + var stalled = watched is not null && !watched.Stop.IsCancellationRequested && inProgress.Any(checkpoint => + checkpoint.CategoryId == watched.CategoryId + && timeProvider.GetUtcNow().UtcDateTime + - (checkpoint.LastProgressAt is { } lastProgress && lastProgress > watched.RunStartedAt ? lastProgress : watched.RunStartedAt) > StallLimit); + + if (stalled) + { + logger.LogError( + "{CategoryId} has committed nothing for {StallLimit}, so it is being stopped and settled Halted. The categories after it still get their go.", + watched.CategoryId, StallLimit); + await watched.Stop.CancelAsync(); + } + } + // The copy finished or the host is stopping: the outer catch ends the poll. + catch (OperationCanceledException) when (cancellationToken.IsCancellationRequested) + { + throw; + } + // Log, don't throw: an error here would come out of DisposeAsync and hide what the copy itself failed with. + catch (Exception exception) + { + logger.LogError(exception, "The stall watchdog's poll failed, so a stalled copy will not be noticed until a later poll succeeds. The copy itself is unaffected and is still running."); + } + } + } + catch (OperationCanceledException) when (cancellationToken.IsCancellationRequested) + { + // Either the copy finished and disposal cancelled the poll, or the host is shutting down. + } + } + + public async ValueTask DisposeAsync() + { + await cancellation.CancelAsync(); + + try + { + await polling; + } + finally + { + cancellation.Dispose(); + } + } + + sealed record RunningCategory(string CategoryId, DateTime RunStartedAt, CancellationTokenSource Stop); + } +} diff --git a/src/ServiceControl/Migration/MigrationStartupCheckRunner.cs b/src/ServiceControl/Migration/MigrationStartupCheckRunner.cs new file mode 100644 index 0000000000..4ccc0626a5 --- /dev/null +++ b/src/ServiceControl/Migration/MigrationStartupCheckRunner.cs @@ -0,0 +1,39 @@ +namespace ServiceControl.Migration; + +using System; +using System.Collections.Generic; +using System.Threading; +using System.Threading.Tasks; +using ServiceControl.Persistence.DataMigration; + +/// +/// Runs startup checks in order and stops at the first one that refuses. +/// +static class MigrationStartupCheckRunner +{ + /// + /// Runs each check in turn, and wraps whatever a failing one throws in a message naming the check and + /// saying that this start has copied nothing, because every check runs before this start's copy. + /// + /// A check failed. The check's own message is kept, and its exception is the inner one. + public static async Task Run(IReadOnlyList checks, CancellationToken cancellationToken = default) + { + foreach (var check in checks) + { + try + { + await check.Run(cancellationToken); + } + // A shutdown is not a check failing, and saying it was would send the customer after the wrong thing. + catch (OperationCanceledException) when (cancellationToken.IsCancellationRequested) + { + throw; + } + catch (Exception exception) + { + throw new Exception( + $"Migration startup check '{check.Name}' failed, so ServiceControl will not start, and this start has copied nothing. {exception.Message}", exception); + } + } + } +} diff --git a/src/ServiceControl/Migration/RecordHostOpenedOnTarget.cs b/src/ServiceControl/Migration/RecordHostOpenedOnTarget.cs new file mode 100644 index 0000000000..676eb714fa --- /dev/null +++ b/src/ServiceControl/Migration/RecordHostOpenedOnTarget.cs @@ -0,0 +1,40 @@ +namespace ServiceControl.Migration; + +using System; +using System.Threading; +using System.Threading.Tasks; +using Microsoft.Extensions.DependencyInjection; +using Microsoft.Extensions.Hosting; +using ServiceControl.Persistence.DataMigration; + +/// +/// Calls when a host started with the migration on opens +/// on a database a migration has already written to. +/// +sealed class RecordHostOpenedOnTarget(IServiceProvider services) : IHostedLifecycleService +{ + // StartedAsync, not StartAsync: the web server binds its port during StartAsync, and a start that dies there + // served nothing. The stores are resolved here rather than injected, because a RavenDB host registers neither + // and is refused before it gets this far. + public async Task StartedAsync(CancellationToken cancellationToken = default) + { + var checkpointStore = services.GetRequiredService(); + + // Only a database a migration has written to gets the stamp, so the marker answers whether a host + // has opened since the copy began rather than whether one ever ran on this database. + if ((await checkpointStore.ReadAll(cancellationToken)).Count > 0) + { + await services.GetRequiredService().RecordHostOpened(cancellationToken); + } + } + + public Task StartingAsync(CancellationToken cancellationToken = default) => Task.CompletedTask; + + public Task StartAsync(CancellationToken cancellationToken = default) => Task.CompletedTask; + + public Task StoppingAsync(CancellationToken cancellationToken = default) => Task.CompletedTask; + + public Task StopAsync(CancellationToken cancellationToken = default) => Task.CompletedTask; + + public Task StoppedAsync(CancellationToken cancellationToken = default) => Task.CompletedTask; +} diff --git a/src/ServiceControl/Migration/RequiredCopyBeforeTheHostOpens.cs b/src/ServiceControl/Migration/RequiredCopyBeforeTheHostOpens.cs new file mode 100644 index 0000000000..9ca9388246 --- /dev/null +++ b/src/ServiceControl/Migration/RequiredCopyBeforeTheHostOpens.cs @@ -0,0 +1,30 @@ +namespace ServiceControl.Migration; + +using System; +using System.Threading; +using System.Threading.Tasks; +using Microsoft.Extensions.Hosting; +using ServiceBus.Management.Infrastructure.Settings; + +/// +/// Runs the required copy as the host starts, before any hosted service of its own does. +/// A refusal throws, which fails the start and stops the host, as the copy requires. +/// +sealed class RequiredCopyBeforeTheHostOpens(IServiceProvider services, Settings settings) : IHostedLifecycleService +{ + // Inside the host rather than before it, because a Windows service reports itself started only once the host + // starts, and the Service Control Manager kills a process that has said nothing for 30 seconds. Every + // StartingAsync runs before any hosted service, so nothing has bound a port or begun ingesting. + public Task StartingAsync(CancellationToken cancellationToken = default) => + MigrationStartup.RunRequiredCopy(services, settings, cancellationToken); + + public Task StartAsync(CancellationToken cancellationToken = default) => Task.CompletedTask; + + public Task StartedAsync(CancellationToken cancellationToken = default) => Task.CompletedTask; + + public Task StoppingAsync(CancellationToken cancellationToken = default) => Task.CompletedTask; + + public Task StopAsync(CancellationToken cancellationToken = default) => Task.CompletedTask; + + public Task StoppedAsync(CancellationToken cancellationToken = default) => Task.CompletedTask; +} diff --git a/src/ServiceControl/Persistence/PersistenceFactory.cs b/src/ServiceControl/Persistence/PersistenceFactory.cs index 59981ba691..5fd045058d 100644 --- a/src/ServiceControl/Persistence/PersistenceFactory.cs +++ b/src/ServiceControl/Persistence/PersistenceFactory.cs @@ -10,6 +10,12 @@ namespace ServiceControl.Persistence static class PersistenceFactory { + /// + /// The manifest names of the two persisters built on EF Core. Anything that is true of both of them and + /// of neither RavenDB nor a future persister is decided by this list. + /// + public static readonly string[] SqlPersistenceNames = ["SQLServer", "PostgreSQL"]; + public static IPersistence Create(Settings settings, bool maintenanceMode = false) { var persistenceConfiguration = CreatePersistenceConfiguration(settings.PersistenceType, settings); diff --git a/src/ServiceControl/Program.cs b/src/ServiceControl/Program.cs index b74bdc8488..96fb2359df 100644 --- a/src/ServiceControl/Program.cs +++ b/src/ServiceControl/Program.cs @@ -52,7 +52,13 @@ LoggingConfigurator.ConfigureNLog("bootstrap.txt", "./", NLog.LogLevel.Fatal); NLog.LogManager.GetCurrentClassLogger().Fatal(ex, "Unrecoverable error"); } - throw; + + // The message goes to the console on its own, because the log above already holds the whole exception. + // Rethrowing instead would print the stack trace a second time and fire the unhandled-exception handler + // for a third, which buries a configuration mistake in what reads like a crash. + await Console.Error.WriteLineAsync(ex.Message); + + return 1; } finally {