Add a D/R mode to the mongo-processor - #2833
Conversation
Hello delthas,My role is to assist you with the merge of this Available options
Available commands
Status report is not available. |
Codecov Report❌ Patch coverage is
Additional details and impacted files
... and 1 file with indirect coverage changes
@@ Coverage Diff @@
## development/9.6 #2833 +/- ##
===================================================
+ Coverage 76.52% 76.76% +0.24%
===================================================
Files 207 211 +4
Lines 14487 14557 +70
===================================================
+ Hits 11086 11175 +89
+ Misses 3391 3372 -19
Partials 10 10
Flags with carried forward coverage won't be shown. Click here to find out more. 🚀 New features to boost your workflow:
|
Waiting for approvalThe following approvals are needed before I can proceed with the merge:
|
1a74a03 to
8675518
Compare
Waiting for approvalThe following approvals are needed before I can proceed with the merge:
|
e77010b to
5112a3a
Compare
5112a3a to
0eb82b1
Compare
2956946 to
d161e55
Compare
d161e55 to
486c8e2
Compare
486c8e2 to
8b693e6
Compare
8b693e6 to
e1dd485
Compare
|
Requested @maeldonn in place of Sylvain Senechal, who is currently on PTO. |
|
Requested @SylvainSenechal in place of Mael Donnart, who is currently on PTO. |
The D/R pipeline should not send every document anymore. Since mongo-processor handles the master "creation/update", the D/R pipeline should now drop the master whenever they have an associated version key (like ingestion populator does). |
df647fc to
9ed5c73
Compare
Done in scality/zenko-operator#631: the pipeline keeps a master only when it is the object itself, i.e. with no version of its own, or a suspended bucket's null version. |
9ed5c73 to
e923570
Compare
Waiting for approvalThe following approvals are needed before I can proceed with the merge:
|
| _updateObjectDataStoreName(entry, location) { | ||
| entry.setDataStoreName(location); | ||
| } | ||
| _getTargetMetadata(log, entry, target, done) { |
There was a problem hiding this comment.
Async/await migration suggestion: _getTargetMetadata is a new function with a done callback parameter. Per the project's async/await migration rule, new functions should use async/await. It could return { objMD, versionId } (or throw on error), and callers could await or util.callbackify it. Not a blocker given the surrounding callback context.
e923570 to
6d7bcb2
Compare
|
Added a |
06718c5 to
aab6f69
Compare
| return false; | ||
| } | ||
|
|
||
| targetVersionId(entry) { |
There was a problem hiding this comment.
note for reviewers: the content of this function could/should probably be moved out of the policy; however
- we don't want to change the behavior of OOB here, so keeping the exact behavior of ingestion policy for now
- ingestion requires the x-scal-version-id mapping anyway, which does not apply for pull replication
Waiting for approvalThe following approvals are needed before I can proceed with the merge:
|
The `server` section is optional, but the whitelist of addresses allowed to reach the health checks was extended unconditionally, so a process serving no API exited before starting: a D/R sink runs the mongo-processor alone, and configuring a section it never reads to get past this is no answer. Issue: BB-811
The mongo-processor reads the object it is about to write, and a first write finds nothing there. Ingestion reaches that path only for a restored object or a bucket carrying replication rules, so both the error log and the error test it leans on went unnoticed; a D/R sink reads the stored document for every entry, which makes a miss the steady state. Log it as the outcome it is, and test it the way the delete path a few lines below already does. `err.NoSuchKey` is arsenal's deprecated comparison, set only while `allowUnsafeErrComp` is on and slated for removal with ARSN-176 -- with it off, every first delivery of every object would be an error. Issue: BB-811
The bucket processor circuit breaker expands a `${location}` template over
every location in the config, and its spec restated the expansion by hand,
one entry per location: adding a location to the sample config broke it.
Build the expectation from the config instead.
Issue: BB-811
The mongo-processor was written as the out-of-band ingestion consumer, and the D/R metadata sink is to reuse it. The two disagree about most of what it does to an object's metadata, so put those decisions behind a policy: an abstract MetadataPolicy whose methods assert, an implementation per mode, and an index mapping the configured mode to the class, as the notification extension does for its destinations. The processor hands the policy entries and asks it whether an entry needs the metadata already stored, which version on this site the entry acts on, how the entry applies onto what is stored -- including the version the result is written under, or nothing if the entry changes nothing -- and whether a delete still applies. IngestionMetadataPolicy carries today's behaviour verbatim, so this commit changes nothing. The `x-amz-meta-scal-version-id` header of a restored object is ingestion's alone, so the policy reads it, where the processor did. The version an entry is written under stays apart from the one it is read from: an entry naming a version that is not stored is ingested under its own id. The default lives beside the policy map, so a processor built programmatically gets the same mode as one built from a config file. Issue: BB-811
The D/R metadata sink replicates production's objects, accounts included, so it applies what the source-side pipeline sends rather than rewriting it into something local: a new object is written as it arrives, with its ACLs reset as they are not replicated, and an update merges the entry's tags and object-lock state into the stored document, cleared values included -- removing a legal hold is an update. Every update is written, a replay included, rather than diffed against what is stored. The stored document keeps everything else, and above all its placement. The copy engine rewrites location and dataStoreName to a local location after the first write, and applying the entry's would send reads back to the source and leak a local copy that is never garbage-collected; a version still on the remote site takes the entry's. A restore this site performed is kept the same way. Several things follow that were unreachable before: - the stored document is always read, because it is what tells a first write from an update, and the entry cannot -- an insert is redelivered on replay and overlaps the bootstrap dump, so it is no promise that the object is absent here; - a delete always applies, where the ingestion guard skips one whose object has moved location, which for a replicated object it always has; - a scal version id names a version of the system an object was ingested from, and is ignored. The D/R source keys every entry by the document it comes from, so the sink acts on that very document: a version by its version id, null versions included, and a master as the master document itself, which only reaches the sink for an object with no version of its own. A master document that copies a version, or the latest version mongo returns in place of a missing master, is not what such an entry targets: the entry is written over it. The mock client returns the document it was given, so null versions are tested against mongodb. Issue: BB-811
An object with no version of its own is rewritten in place, and so is the master a versioning suspended bucket marks null. Merging such an entry kept the previous object's content, headers, user metadata and archive on the sink, which then described an object that no longer existed. The Kafka Connect sink this replaces wrote the whole document, so the merge was a regression against it. An overwrite moves the modification date, which a tag, retention, legal hold or restore update keeps: an entry whose date differs from the stored one replaces the document, and any other entry still merges, so a tag change keeps a restore this site performed. The content digest cannot tell an overwrite either, a PUT replacing the whole metadata whatever the bytes do. A version is immutable and always merges. Nothing of a replaced document is kept. A restore this site performed describes bytes that are gone; reclaiming them is left to the D/R garbage collection still to come. Issue: BB-811
Every backbeat process reads its mongodb client from the queuePopulator section, so a process that only needs that client, like the mongo-processor of a D/R sink, had to fill in a cron rule, a zookeeper path, a probe server and a log source for a populator that never runs. The 'none' log source says so: it needs none of those settings, and a queue populator started with it refuses to run. Issue: BB-811
A null master is a version like any other: the source gives it a document of its own on its first metadata update, and streams every later change under that version id. Requesting and writing it as the master would leave those changes on another document, so the sink targets a null master by the internal version id it carries, and only an object with no version id at all is still written in place, told from an overwrite by its modification date. That also leaves one place to decide whether an entry updates the stored document: a targeted version always does, and the guard for a master copying a version has no case left to cover. Issue: BB-811
aab6f69 to
d558470
Compare
|
/bypass_author_approval |
|
I have successfully merged the changeset of this pull request
The following branches have NOT changed:
This pull request did not target the following hotfix branch(es) so they
Please check the status of the associated issue BB-811. Goodbye delthas. The following options are set: bypass_author_approval |
The D/R metadata sink reuses the mongo-processor (ZKOP-562). Ingestion and D/R disagree about what to do with an object's metadata, so those decisions move behind a
MetadataPolicy, picked byextensions.mongoProcessor.mode:ingestion, the default, carries today's behaviour verbatim;drselectsPullReplicationMetadataPolicy.The processor calls the policy hooks, drawn as hexagons, at these points:
flowchart TD E["entry from Kafka"] --> T{"put or delete?"} T -->|put| S1{{"skipsMetadataFetch"}} S1 -->|false| S2{{"targetVersionId"}} S2 --> R1["read the stored document"] S1 -->|true| A{{"apply"}} R1 --> A A -->|null| K1["skip"] A -->|"content, versionId"| RI["resolve replicationInfo<br/>against this site's bucket"] RI --> W["write under versionId and repair the master,<br/>or write the master alone"] T -->|delete| D1{{"targetVersionId"}} D1 --> R2["read the stored document"] R2 -->|missing| K2["skip"] R2 --> D2{{"skipsDelete"}} D2 -->|true| K2 D2 -->|false| DEL["delete the version and repair the master,<br/>or delete the master alone"]"Overwritten in place" is an object with no version of its own, or the master a suspended bucket marks null. An overwrite moves the modification date, which a tag, retention, legal hold or restore update keeps.
microVersionIdcannot tell an overwrite from an update: tagging, retention, legal hold, ACL and restore completion bump it, and all of those merge, while a plain PUT never sets it.The stored placement is kept unless the stored location is still on the remote site (
isCRR), which is groundwork for clean room. On a deployed PRA sinklocationConfig.jsonis{}, so that branch never fires.Commits, in order: no health-check whitelist without a
serversection, as a D/R sink serves no API; two small fixes; the policy, which moves ingestion's behaviour unchanged; the D/R policy; then the overwrite rule; and anonelog source for a process that runs no queue populator, so the sink'squeuePopulatorsection needs nothing beside its mongodb client. The sink reads its mongodb client fromqueuePopulator.mongoand its service account fromextensions.ingestion.auth, as the rest of backbeat does; zenko-operator#631 writes those sections.Known limits:
replicationInfois reset on every write: the source pipeline strips the replication configuration from sink buckets.Follow-ups: data GC on the sink (BB-813, BB-815), circuit breaking (BB-822),
__metastoreand Vault entities (OS-1110),transitionInProgress(BB-819), a reallocationConfig.jsonon the sink (ZKOP-566).Goes with scality/zenko-operator#631, which keys every entry by the document it comes from.
Issue: BB-811