Skip to content

Fix witness - #96

Merged
On1x merged 406 commits into
masterfrom
fix-witness
May 17, 2026
Merged

On1x merged 406 commits into
masterfrom
fix-witness

Conversation

@On1x

@On1x On1x commented Apr 28, 2026

Copy link
Copy Markdown
Member

No description provided.

On1x added 30 commits May 3, 2026 18:19
- Add transaction count to the block generation log message for better visibility
- Include the number of transactions in the debug capture when broadcasting block
- Enhance debug information during block production with transaction details
…point

- Added logic to determine the originating peer endpoint from active connections
- Passed the identified peer endpoint to handle_block instead of an empty optional
- Enhanced tracking of items being processed to improve block acceptance handling
- Set we_need_sync_items_from_peer to false to fix synchronization logic
- Update last stale block check position upon receiving a new block in p2p plugin
- Improve tracking of last stale head block for sync status detection
…entation

- Detail full activation process and guards in database update_global_dynamic_data
- Explain emergency mode deactivation criteria and hybrid witness schedule override
- Describe witness block production behavior and master/follower roles in emergency
- Document fork database deterministic tie-breaking for emergency block conflicts
- Clarify last irreversible block advancement caps during emergency mode
- Outline startup recovery steps to repair emergency witness schedule
- Describe P2P plugin guards for stale sync detection and emergency resync logic
- Explain witness guard key restoration behavior specific to emergency mode
- Include snapshot plugin handling of emergency state fields during import
- Provide full system state diagram and component interaction mapping
- Summarize all emergency mode guards, safety invariants, and versioning exclusions
…spam)

When a DLT slave processes a burst of sync blocks (e.g. 21 blocks in
80ms), two reinforcing message storms stall block reception from the
master:

Sending side: send_sync_block_to_node_delegate() notifies EACH in-sync
peer about EACH accepted block via fetch_next_batch_of_item_ids_from_peer(),
generating a full get_blockchain_synopsis() per call. This creates
O(N*M) synopsis computations that flood the event loop.

Receiving side: When multiple peers send fetch_blockchain_item_ids_message
requests simultaneously (triggered by our outbound notifications), we
respond to each by calling the expensive get_block_ids(), generating
redundant computations.

Fix sending side: Add per-peer 5-second cooldown on in-sync
notifications in send_sync_block_to_node_delegate(). Also skip if a
synopsis request is already pending for the peer (item_ids_requested_from_peer).

Fix receiving side: Add per-peer 5-second rate limit on get_block_ids()
responses in on_fetch_blockchain_item_ids_message() for in-sync peers.
Critical: only skip if our head block number hasn't changed since the
last response -- if we've accepted new blocks, the peer needs updated
IDs and we must respond normally even within the cooldown.

New fields in peer_connection:
- last_in_sync_notification_sent: cooldown guard for outbound notifications
- last_fetch_item_ids_response_time: rate-limit guard for inbound responses
- last_fetch_item_ids_response_head_num: prevents rate-limit from hiding
  new blocks when head has advanced since last response
…d guard

Receiver: increment strikes on rate-limited in-sync fetch_item_ids requests;
soft-ban peer after 50 strikes (5min). Demote per-hit log to dlog.

Sender: add pending-request guard + 100ms cooldown to
fetch_next_batch_of_item_ids_from_peer() to prevent outbound synopsis loops.
… all of a peer's synopsis entries are below our DLT range. Instead of throwing peer_is_on_an_unreachable_fork, it uses the highest synopsis entry as an anchor and sets found_a_block_in_synopsis = true.
…y consensus

- Define new chain_status_announcement_message with relevant chain and consensus fields
- Add core_message_type_enum entry for chain_status_announcement_message_type (5018)
- Extend node interface with DLT and emergency consensus related query methods
- Track peer DLT and emergency consensus info during handshake in peer_connection
- Include DLT and emergency consensus info in user_data sent during hello message exchange
- Implement p2p_plugin methods to query and broadcast chain status announcements on connection changes
- Handle receiving chain_status_announcement_message in p2p_plugin and log details
- Broadcast chain status announcement upon new peer connection establishment
- Use weak read locks when accessing database state for chain status data
- Implement heuristic to identify emergency key holder using block production status
… updates

- Introduce message type 5018 to broadcast live chain state including head block, LIB,
  DLT mode window, and emergency consensus status to peers
- Extend hello message with DLT and emergency consensus related fields for initial snapshot
- Implement handler for chain_status_announcement_message to log and update peer's live status
- Broadcast chain status automatically on new peer connections and allow manual broadcast
- Prevent stale DLT window and emergency consensus related sync issues and fork exceptions
- Maintain backward compatibility with peers lacking support for new message type
- Provide detailed documentation outlining benefits, protocol design, and integration with existing protections
….hpp:94. This is safe because core_message_type_last is only used as an upper-bound sentinel in a range check
… loop

- Add bounds checking in appbase write_default_config() to prevent
  std::out_of_range when format_parameter() returns short strings.
- Skip db.revision() check in replay_db() when force_replay is already
  true, avoiding a secondary boost::interprocess::lock_exception on
  corrupted shared memory during recovery.
database::open(): catch boost::interprocess::lock_exception during
undo_all() and rethrow as database_revision_exception so the chain
plugin's snapshot-recovery path handles it instead of std::terminate.

update_global_dynamic_data(): don't reset current_run for the
block-producing witness during missed-slot penalty processing.
In emergency hybrid mode, the committee witness occupies multiple
schedule slots; resetting current_run for its own duplicate slots
prevented it from reaching CHAIN_IRREVERSIBLE_SUPPORT_MIN_RUN,
stalling LIB advancement indefinitely.

Update appdata submodule.
- Removed unnecessary const references when calling _db.get_account() in multiple evaluator files
- Deprecated commented assertions related to hardfork 9 left in place for future removal
- Added new network node method declarations for DLT mode and emergency consensus features
- Improved code readability by eliminating unused variable bindings in account fetching operations
… internal wdump(__VA_ARGS__) fails with zero arguments. Since there are no variables to log, FC_CAPTURE_AND_RETHROW() is the correct macro (same rethrow behavior without the warning dump).
The DLT sync guard (db._dlt_mode && chain().is_syncing()) in maybe_produce_block() blocked block production for emergency
consensus master nodes after restart. The master is the sole block producer — waiting for sync to complete creates a permanent
deadlock because no blocks arrive to clear the syncing flag.

Fix: add !dgp.emergency_consensus_active to the guard condition, allowing the emergency master to produce blocks regardless of
sync state. The oscillation fix for non-emergency DLT mode is preserved.
The snapshot plugin's stalled sync detection did not distinguish
between the emergency master and follower nodes. When the master
produces solo blocks during emergency consensus (normal behavior
since it IS the chain), the stalled sync timer would expire and
trigger unnecessary recovery, potentially disrupting the network.

Add witness_plugin::is_emergency_master() implementing the dual
identity check: (1) committee account is in _witnesses (proving
emergency-private-key is configured), AND (2) committee appears
in the current witness schedule. Use this in the snapshot plugin
to skip recovery entirely when we are the master  only followers
on isolated forks need recovery.
…from competing forks in DLT+emergency mode. Blocks at/below head with different ID, or blocks above head whose parent doesn't match our head, are skipped before entering accept_block/push_block — avoiding the unlinkable_block_exception → DEFERRED_RESIZE cascade that starves block production and crashes the node.
…on — now the condition only checks:

block_num > head_block_num
previous != block_id_type() (genesis)
previous != head_block_id
On1x added 29 commits May 14, 2026 15:43
…ync task

- Add flag_guard struct to manage snapshot_in_progress flag with explicit release method
- Release flag and resume P2P processing after DB read lock drops in snapshot callback
- Log message confirms DB read completion and background compression/file writing
- Safely resume P2P if create_snapshot throws before callback execution
- Use memory_order_release for atomic flag updates to ensure proper synchronization
- Add catch block for shared_memory_corruption_exception during block acceptance
- Log the corruption details with block number
- Trigger chain auto-recovery to prevent node from getting stuck
- Return rejected result after initiating recovery to retry processing
- Prevent exception from falling through to generic handler causing silent rejection
- Include p2p_plugin in chain plugin dependencies and headers
- Pause all P2P database consumers before closing the corrupted database to prevent data corruption
- Mark syncing state to defer witness block production during recovery
- Set snapshot path and trigger full database recovery (wipe, import, DLT replay)
- Resume P2P block processing after the database is fully rebuilt
- Add error handling and logging for pause/resume operations on P2P plugin
- Added PRIVATE include directory for p2p plugin headers in chain plugin
- Updated target_include_directories with new path in CMakeLists.txt
- Ensured proper separation of PUBLIC and PRIVATE include paths
- Adjusted CMakeLists.txt formatting for consistency
- Replace deprecated graphene_p2p_plugin with graphene::p2p
- Adjust CMake target dependencies accordingly
- Maintain target include directories without changes
- Extract set_resume_flags() to atomically update resume flags from any thread
- Separate run_resume_on_p2p_thread() for logging and draining on P2P thread
- Modify resume_block_processing() to call set_resume_flags() and run_resume_on_p2p_thread()
- Update p2p_plugin to set flags before async drain without waiting on P2P thread
- Ensure _paused_block_queue draining only occurs on P2P thread to avoid threading issues
- Clean up comments for better clarity on threading and flag usage
- Add check to identify if block producer is one of our own witnesses
- Only count slot hijack if producer is not one of our witnesses
- Reset hijack counter correctly when our witness produces the block
- Update logging to reflect actual producing witness in hijack resolution message
- Ensure _catchup_after_pause is set before clearing _block_processing_paused
- Prevent witness thread from reading inconsistent resume flags state
- Add early return in witness plugin when no scheduled witnesses exist
- Avoid unnecessary processing when num_witnesses is zero
…back

- Update comment to clarify current_aslot usage after apply_block call
- Fix slot index calculation by removing minus one offset
- Ensure correct identification of scheduled witness slot during callback
When a witness misses a slot due to timing lag (>500ms), the production
loop was entering a tight loop rechecking the same missed slot every
250ms. Since the head block hadn't advanced, get_slot_at_time() kept
returning the same slot number, causing infinite lag returns.

Added skip-ahead logic to track the lagged slot time and jump past the
missed slot interval, preventing the tight loop and allowing production
to wait for the next actual slot boundary.

Changes:
- Add _last_lag_slot_time member to track last lag condition
- Store scheduled_time when lag is detected
- Skip ahead past missed slot interval before rescheduling
- Return early to avoid double-scheduling the production timer
…kip units

Root cause (p74): production_timer_ shared appbase's global io_service with
the P2P layer. When peer 138.201.117.201 was disconnected (send queue=100,
ECANCELED + drain_send_queue callbacks), the io_service backlog delayed the
production timer callback by 1262ms, causing a missed slot for dota2time.

Move production_timer_ to a dedicated asio::io_service with its own thread
so P2P network activity can never delay the production tick. Shutdown via
io_service::stop() which is thread-safe; destructor guards against missing
plugin_shutdown().

Also fix a units bug in the lag skip-ahead logic: time_since_lag was computed
in seconds (count()/1000000) but subtracted from a millisecond value, yielding
skip_ms = 2999 instead of ~1500 — causing the loop to overshoot past the next
slot. Now computed in ms throughout.
- Replace bool block-processing flags with std::atomic<bool> for thread-safe pause/resume
- Refactor resume_block_processing() into two-phase design to avoid fiber deadlock
- Fix atomic flag ordering to prevent stale-head fork detection issues
- Handle shared_memory_corruption_exception in block acceptance to trigger auto-recovery
- Clear _dlt_syncing flag on SYNC to FORWARD transition to unblock witness production
- Add node uptime field to DLT Status and P2P Stats log lines

feat(witness): enhance production timer and add watchdog diagnostics

- Run production timer on dedicated io_service and thread to avoid P2P I/O delays
- Prevent lag tight loop by skipping ahead after missed slot detection
- Add production watchdog firing on no block produced for 60s or 180s based on witness role
- Implement slot hijack detection counting consecutive emergency master filled slots
- Fix false-positive hijack counter reset when own-witness blocks produced
- Add not_my_turn streak detection warning after prolonged schedule misalignment
- Add missed-slot diagnostic via on_block_applied signal dumping plugin state
- Refine slot=0 stall detection to count only real NTP stalls, not normal waits

fix(snapshot): correct snapshot plugin initialization and async resume flags

- Fix snapshot path propagation into async load callback to avoid empty paths
- Fix initialization order in snapshot plugin startup to prevent uninitialized state access
- Fix P2P resume flags (_block_processing_paused and catchup) reset reliably after async snapshot task completes on non-P2P thread
- Document two distinct periodic log lines emitted by dlt_p2p_node
- Explain fields and usage of the compact DLT Status line emitted every ~30 seconds
- Provide examples and interpretation guidelines for uptime and flags
- Add corresponding Russian translation with parity to English version

fix(p2p_plugin): resolve deadlock in resume_block_processing with async wait

- Split resume_block_processing into atomic flag setter and async post phases
- Remove blocking async().wait() call to avoid P2P thread deadlock under read lock
- Ensure correct ordering of _catchup_after_pause and snapshot_in_progress flags
- Update flags to std::atomic<bool> to support lock-free thread-safe transitions

fix(snapshot-plugin): ensure flags reset after snapshot async task failure

- Add RAII guard to call resume_block_processing on all async task exit paths
- Pass snapshot path correctly into async load callback to avoid silent load failure

refactor(witness-plugin): isolate production loop on dedicated io_service thread

- Move production loop timer to dedicated production_io_service_ and thread
- Prevent P2P network activity from delaying production timer callbacks
- Add lag skip guard to avoid spinning on missed slots by skipping ahead
- Improve slot hijack detection logging and reset logic under emergency mode
- Introduce _last_lag_slot_time and _watchdog_debug_enabled state variables for diagnostics
… fix p2p broadcast

──────────────────────────────────────────
1. Hardfork 13: Validator Reward Sharing
──────────────────────────────────────────
Validators can now redirect a fraction of their block reward to the accounts
that voted for them (stakeholders), accumulated as TOKEN and converted to
SHARES at each distribution epoch end.

New chain property (consensus): distribution_epoch_length
  Validators vote via chain_properties_hf13 struct.  Median across
  scheduled validators becomes the consensus value.  Default: 28800 blocks
  (1 day).  Range: [CHAIN_MIN_DISTRIBUTION_EPOCH_LENGTH=21,
  CHAIN_BLOCKS_PER_YEAR].

New operation: set_reward_sharing_operation
  Requires active authority of owner.  Field sharing_rate is in basis
  points (0 = no sharing, 10000 = 100%).  Rejected before HF13.
  Validator must have a registered witness_object.

Block reward split (process_funds):
  stakeholder_token = witness_reward * sharing_rate / CHAIN_100_PERCENT
  validator_token   = witness_reward - stakeholder_token
  Validator receives SHARES immediately via create_vesting().
  stakeholder_token is accumulated in witness_object.pending_stakeholder_reward
  as TOKEN (already counted in current_supply; not yet in total_vesting_fund).
  When sharing_rate == 0, full reward goes to validator unchanged.

Epoch distribution (process_validator_epoch_distribution):
  Called at end of _apply_block when head_block_num % distribution_epoch_length == 0.
  Uses time-weighted per-stakeholder shares to prevent flash-voting:
    witness_vote_object gains vote_created_block field (set on vote creation;
    zero for pre-HF13 votes, treated as full-epoch weight).
    weighted = witness_vote_weight() * (head_block_num - max(vote_created_block,
                                        epoch_start_block) + 1)
  Stakeholders with computed share < CHAIN_MIN_STAKEHOLDER_REWARD_PAYOUT (1 atomic
  TOKEN = 0.001 VIZ) are skipped; their portion reverts to the validator as dust.
  Dust is NOT burned — reward sharing is the validator's voluntary decision; a
  stakeholder who fails to meet the minimum threshold has simply not earned a
  viable payout.  No-stakeholder case: full pool returned to validator.
  Emits stakeholder_reward_operation per eligible stakeholder.

New virtual operation: stakeholder_reward_operation
  Fields: validator (account_name_type), stakeholder (account_name_type),
          shares (asset, SHARES credited to stakeholder).
  Registered in account_history plugin for both validator and stakeholder.

Chain object changes (witness_objects.hpp):
  witness_object: + sharing_rate (uint16_t), + pending_stakeholder_reward (share_type)
  witness_vote_object: + vote_created_block (uint32_t)

New constants (config.hpp):
  CHAIN_VALIDATOR_MAX_SHARING_RATE = CHAIN_100_PERCENT
  CHAIN_MIN_STAKEHOLDER_REWARD_PAYOUT = int64_t(1)
  CHAIN_DEFAULT_DISTRIBUTION_EPOCH_LENGTH = CHAIN_BLOCKS_PER_DAY
  CHAIN_MIN_DISTRIBUTION_EPOCH_LENGTH = uint32_t(CHAIN_MAX_WITNESSES)
  CHAIN_VERSION bumped 3.1.0 → 3.2.0

apply_hardfork(13): no migration needed; sharing_rate and
  pending_stakeholder_reward default to 0 (initialised on replay).

Files: hardfork.d/0-preamble.hf, hardfork.d/13.hf, config.hpp,
  chain_operations.hpp, chain_virtual_operations.hpp, operations.hpp,
  chain_operations.cpp, chain_evaluator.hpp, chain_evaluator.cpp,
  chain_properties_evaluators.cpp, witness_objects.hpp, database.hpp,
  database.cpp, account_history/plugin.cpp

──────────────────────────────────────────
2. Witness → Validator CLI and config rename
──────────────────────────────────────────
Production node options are renamed from witness-centric to validator-centric
language.  Deprecated aliases are preserved so existing config.ini files
continue to work without modification; they emit a wlog() deprecation warning.

witness.hpp / witness.cpp:
  --witness / -w  → --validator / -v  (deprecated alias kept)
  production_condition enum values: renamed to validator_* prefix
  Internal log strings updated to "validator"

witness_guard.cpp:
  --witness-guard-enabled    → --validator-guard-enabled
  --witness-guard-witness    → --validator-guard-validator
  --witness-guard-interval   → --validator-guard-interval
  --witness-guard-disable    → --validator-guard-disable
  Priority: new option wins; fall back to deprecated option if only it is set.

config_witness.ini:
  Updated example config to use new option names; deprecated names kept as
  commented-out examples with DEPRECATED markers.

──────────────────────────────────────────
3. Fix: decouple block post-validation broadcast from production timer
──────────────────────────────────────────
The production timer thread called p2p().broadcast_block_post_validation()
which internally does p2p_thread.async(...).wait() — blocking the production
thread until the P2P thread finishes writing to ALL peers, including slow ones.

Root cause: peer 37.192.123.64 accumulated a send queue of ~100 messages over
~3 minutes.  When the production timer callback called
broadcast_block_post_validation() 20 times per tick, each .wait() held the
production thread until the P2P thread completed its socket write to that peer.
Cumulative blocking exceeded the 3-second slot window, causing witness "id" to
miss its slot at block ~#79941785.  After the miss, the witness was voted out of
the schedule, resulting in 252 seconds of silence until the watchdog recovered it.

Add p2p_plugin::post_broadcast_block_post_validation() — a fire-and-forget
variant that posts the task onto the P2P thread via p2p_thread.async() without
.wait().  The production timer thread returns immediately and is never blocked
by peer socket I/O.  The existing blocking variant is preserved for callers
outside the production loop that require delivery confirmation.

witness.cpp: production loop's block_post_validation signing loop now calls
  post_broadcast_block_post_validation() instead of the blocking variant.

Files: p2p_plugin.hpp, p2p_plugin.cpp, witness.cpp

──────────────────────────────────────────
4. Documentation
──────────────────────────────────────────
.qoder/docs/hf13-validator-reward-sharing.md — full HF13 spec: operation
  definitions, epoch distribution algorithm, flash-voter protection analysis,
  token accounting, mid-epoch rate change behaviour, dust rationale.
.qoder/repowiki/en/content/Witness.md — HF13 section added; class diagram
  updated; Update Summary entry added.
- Define HF13 hardfork with future timestamp and increment CHAIN_NUM_HARDFORKS to 13
- Add new operations: set_reward_sharing_operation, stakeholder_reward_operation, and related protocol structs
- Increment CHAIN_SCHEMA_VERSION to 13 due to adding fields to witness_object and witness_vote_object
- Implement proactive schema version check in chain plugin to wipe shared_memory.bin before db.open() if mismatch detected
- Extend account_history plugin to handle new virtual operation for stakeholder rewards
- Update witness_api plugin with preferred validator API method names alongside legacy witness names
- Modify witness_api_object to include sharing_rate and pending_stakeholder_reward fields with zero defaults
- Add deployment guidance for shared memory compatibility and automatic/manual recovery on schema mismatch
- Register HF13 hardfork, evaluator, and new consensus parameters in database.cpp
- Improve schema_version file read/write helpers and integrate version check in plugin startup sequence
- Remove unused helper function for schema version path
- Use Boost filesystem functions directly for file existence checks
- Replace chainbase database wipe with explicit shared memory file removal
- Improve consistency in file path usage for schema version read/write methods
- Add warning log after shared memory file removal to indicate recovery path
- Update schema version check comments for clarity on binary incompatibility
- Remove shared_memory.bin file when schema versions differ
- Write new schema version before recovery to prevent repeated wipes on crash
- Add explicit local snapshot recovery after schema wipe if auto recovery enabled
- Log snapshot import attempts and fallback scenarios during schema wipe recovery
- Ensure normal DB open for P2P sync or replay if no local snapshot available
…on logic

- Add set_reward_sharing method to wallet API for setting validator reward sharing rate
- Include validation and transaction signing in set_reward_sharing operation
- Modify database reward distribution to only create vesting if validator token is greater than zero
- Add documentation comments for set_reward_sharing API method
- Register set_reward_sharing operation in wallet interface functions list
…ote weight

- Clarify that vote weight uses witness_vote_fair_weight() instead of witness_vote_weight()
- Explain fair vote weight as total stake divided by number of validators voted for
- Align documentation with consensus scheduling logic for per-validator weight
- Update code to apply fair vote weight when calculating stakeholder contributions
- Maintain time-weighted distribution description unchanged
- Include sharing_rate and pending_stakeholder_reward for validator reward sharing
- Add vote_created_block for flash-voter protection
- Set default values for these fields in pre-HF13 snapshots to maintain compatibility
…e timing

- Change minimum distribution epoch length to 1 hour to reduce processing frequency
- Add comments explaining the rationale behind increasing minimum epoch length
- Move schema version write to after successful snapshot recovery to prevent skipping retries on failure
- Update comments to clarify recovery and schema version handling during chain plugin open process
- Introduced _diag_mutex to guard diagnostic fields shared between production
  IO thread and P2P thread, preventing data races
- Added lock guards in get_production_diagnostics() and on_block_applied()
- Ensured atomic updates and snapshots of diagnostic counters under mutex
- Wrapped state resets and condition checks with locking to maintain consistency
- Fixed logging statements to use snapshots taken under lock to avoid races
- Preserved _watchdog_debug_enabled access as production-only without lock
- Replace v.contains() with v.get_object().contains() to correctly check JSON keys
- Ensure compatibility with nested JSON objects in snapshot plugin parsing
- Prevent potential crashes or incorrect data extraction during deserialization
- Add catch block for std::exception in dlt_block_log replay
- Log error message and continue with snapshot state on std::exception
- Ensure node proceeds to P2P sync to fill blockchain gaps
- Add conditional logging for missed blocks by non-emergency offline witnesses
- Include witness owner, missed block number, scheduled time, and next witness in log
- Use colored output for better visibility of missed block warnings
- Ensure log only triggers for valid missed witness scenarios to avoid noise
…ssed

- Increase witness scheduling check window from 4 to 5 slots (~15 seconds)
- Add pending_snapshot_safe_after_time to track earliest safe time for deferred snapshot start
- Modify witness scheduling check to store next validator slot time for snapshot deferral
- Defer snapshot creation until head_block_time reaches or exceeds stored safe time
- Log deferred snapshot status when waiting for witness slot to complete
- Introduce get_next_validator_slot_time() in witness plugin to provide slot timing info
- Update snapshot plugin to avoid infinite deferral by checking slot time instead of re-checking witness schedule
- Introduce optional distribution_epoch_length field in chain_api_properties
- Assign distribution_epoch_length from source when hardfork 13 is active
- Update reflective macros to include distribution_epoch_length property
@On1x
On1x merged commit 953b9e0 into master May 17, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant