# Node sync issues on sep-03

**URL:** <https://forum.autonomys.xyz/t/node-sync-issues-on-sep-03/4409>\
**Category:** Support\
**Created:** [September 4, 2024, 7:07am UTC](https://forum.autonomys.xyz/t/node-sync-issues-on-sep-03/4409 "2024-09-04T07:07:31Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![vexr](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/vexr/32/3498_2.png) [@vexr](https://forum.autonomys.xyz/u/vexr)\
**Post date:** [September 4, 2024, 7:07am UTC](https://forum.autonomys.xyz/t/node-sync-issues-on-sep-03/4409/1 "2024-09-04T07:07:31Z")

</div>

I have multiple nodes (6 of them) that are unable to reach the block tip on sep-03 that have no issue on jul-29.

Two are consensus nodes and four are domain nodes.

### Environment

- **Operating System:** Ubuntu 24.04 (VM) Dual Xeon 2680 v3
- **Space Acres/Advanced CLI/Docker:** subspace-node-ubuntu-x86\_64-v2-gemini-3h-2024-sep-03

### Problem

It does sync a few blocks, but is not able to keep up with the network and just falls behind. No issues when reverting to previous release (jul-29).

_I do not have this issue running on subspace-node-ubuntu-x86\_64-skylake-gemini-3h-2024-sep-03 with a non-domain node._

Parameters:  
–sync full  
–blocks-pruning archive-canonical  
–state-pruning archive-canonical

[Log](https://logs.atc.farm/sep-03/node.log)

---

<div class="post-metadata">

**Author:** ![Jim-Autonomys](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/jim-autonomys/32/2922_2.png) [@Jim-Autonomys](https://forum.autonomys.xyz/u/Jim-Autonomys)\
**Post date:** [September 4, 2024, 9:25am UTC](https://forum.autonomys.xyz/t/node-sync-issues-on-sep-03/4409/2 "2024-09-04T09:25:36Z")

</div>

Hey vexr, are you on domain 0 (Nova) on all these operators? I’ve just successfully got to chainhead on domain 1 (auto-id) with `sep-03`.

[Here](https://gist.github.com/jim-counter/2bddaa07b705bb308b8cfc2cec2b7e4c) are some logs if you want to compare. I did spot all the `Failed to dial peer during bootstrapping` at the start but don’t have time right now to go through them properly.

---

<div class="post-metadata">

**Author:** ![vexr](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/vexr/32/3498_2.png) [@vexr](https://forum.autonomys.xyz/u/vexr)\
**Post date:** [September 4, 2024, 9:54am UTC](https://forum.autonomys.xyz/t/node-sync-issues-on-sep-03/4409/3 "2024-09-04T09:54:22Z")

</div>

> [@Jim-Autonomys](#):
>
> are you on domain 0 (Nova) on all these operators? I’ve just successfully got to chainhead on domain 1 (auto-id) with `sep-03`

I am running multiple nodes for different purposes, each running on the same system configuration within its own VM environment. Of those that didn’t reach the chainhead, two were non-domain nodes, two were on domain 0, and two were on domain 1.

Although this issue seems specific to the v2 release in a VM environment, I wanted to share my experience. The log I provided was brief, but I updated almost immediately and dealt with the issue for about four hours before reporting it.

Since my original post, I’ve been running sep-03 continuously on half of the nodes (one on each domain), but the sep-03 builds haven’t been able to keep up, while I’ve had no issues on jul-29.

---

<div class="post-metadata">

**Author:** ![Jim-Autonomys](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/jim-autonomys/32/2922_2.png) [@Jim-Autonomys](https://forum.autonomys.xyz/u/Jim-Autonomys)\
**Post date:** [September 4, 2024, 10:06am UTC](https://forum.autonomys.xyz/t/node-sync-issues-on-sep-03/4409/4 "2024-09-04T10:06:34Z")

</div>

Interesting. I am running the same parameters but am on Skylake bare metal. Thanks for reporting!

@ning any ideas what could be going on here?

---

<div class="post-metadata">

**Author:** ![ning](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/ning/32/669_2.png) [@ning](https://forum.autonomys.xyz/u/ning)\
**Post date:** [September 4, 2024, 3:50pm UTC](https://forum.autonomys.xyz/t/node-sync-issues-on-sep-03/4409/5 "2024-09-04T15:50:20Z")

</div>

Currently, the domain blocks are all derived from the consensus block locally, so if the sync of the consensus chain is slow or paused the domain chain will also progress slow or paused, since there are also consensus node having this issue I think the issue is related to the consensus chain syncing.

From the above consensus node log, after the node is started, it progresses slowly for the first 10 minutes (#3131593 → #3131614) and then the node gets stuck at #3131614 and having error:

```auto
2024-09-04T06:53:31.728618Z WARN Consensus: sync: 💔 Error importing block 0x6baed6897ecbbff758ce476805fc1c7c17dc08189cf34fb99c6c83df0173d56d: block has an unknown parent    
2024-09-04T06:53:31.729262Z WARN Consensus: sync: 💔 Error importing block 0x9e8bcb6d6cdd6ac944f76f7e2a9c264f31555ba6412b95e56bf8e9fedcbeeb19: block has an unknown parent    
2024-09-04T06:53:31.729751Z WARN Consensus: sync: 💔 Error importing block 0x2f6d9a419c515cde0b5588b60a6b8daf7a8ac4be243800c5b167f74d5a7425fa: block has an unknown parent 
...

```

I checked from the gemini-3h consensus RPC node, these error logs point to the blocks following #3131614, i.e. #3131616, #3131617, #3131618…, etc, **but except #3131615** and then the node is synced to `#3131615 (0x38fc…270f)` which is in the stale fork because from the RPC node the hash of #3131615 is 0xe5c6509154fa7ab6295421fec2ade76b71d3305439808e7915f18d86bbbac82e, after that the node stuck at #3131615 till the end of the log.

So I guess there is an issue with the consensus chain networking cause the node can’t fetch the correct #3131615 in the canonical fork, @nazar-pc or @shamil plz take a look.

---

<div class="post-metadata">

**Author:** ![nazar-pc](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/nazar-pc/32/6_2.png) [@nazar-pc](https://forum.autonomys.xyz/u/nazar-pc)\
**Post date:** [September 5, 2024, 12:25pm UTC](https://forum.autonomys.xyz/t/node-sync-issues-on-sep-03/4409/6 "2024-09-05T12:25:57Z")

</div>

This looks suspiciously similar to [Blocks downloaded from peers undergoing reorg are treated as extending canonical chain and fail to import · Issue #2094 · paritytech/polkadot-sdk · GitHub](https://github.com/paritytech/polkadot-sdk/issues/2094) that is a long-standing known issue in Substrate. It should be possible to step over it once archiving processes this block because sync from DSN is not affected by this issue, which is why we have closed [Synced onto a fork · Issue #1744 · autonomys/subspace · GitHub](https://github.com/autonomys/subspace/issues/1744) eventually.

---

<div class="post-metadata">

**Author:** ![vexr](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/vexr/32/3498_2.png) [@vexr](https://forum.autonomys.xyz/u/vexr)\
**Post date:** [September 6, 2024, 6:19am UTC](https://forum.autonomys.xyz/t/node-sync-issues-on-sep-03/4409/7 "2024-09-06T06:19:32Z")

</div>

> [@nazar-pc](#):
>
> It should be possible to step over it once archiving processes this block because sync from DSN is not affected by this issue

I tried running it again (sep-03) and let it go for six hours. The log shows that while it downloads some blocks (albeit slowly), it never fully catches up. However, after reverting to the jul-29 version, it quickly processed all the backlog almost immediately.

Any thoughts on how to resolve?

[Log](https://logs.atc.farm/sep-03/node2.log)

---

<div class="post-metadata">

**Author:** ![nazar-pc](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/nazar-pc/32/6_2.png) [@nazar-pc](https://forum.autonomys.xyz/u/nazar-pc)\
**Post date:** [September 6, 2024, 6:41am UTC](https://forum.autonomys.xyz/t/node-sync-issues-on-sep-03/4409/8 "2024-09-06T06:41:37Z")

</div>

Try to copy the database and run just consensus node with that copy (but with the same sync/pruning options).

We did upgrade Substrate and there could be some upstream or downstream changes affecting this, but the first thing is to identify if domains have anything to do with it.

Also I see you have 128 connections instead of normal 40, removing customization for number of peers might help as well depending on what the root cause of the issue is. There shouldn’t be any reason for you to benefit from it anyway.

---

<div class="post-metadata">

**Author:** ![vexr](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/vexr/32/3498_2.png) [@vexr](https://forum.autonomys.xyz/u/vexr)\
**Post date:** [September 6, 2024, 7:04am UTC](https://forum.autonomys.xyz/t/node-sync-issues-on-sep-03/4409/9 "2024-09-06T07:04:44Z")

</div>

I should have clarified earlier: this latest test was conducted on a consensus node only. I’m encountering the same issue with domain nodes as well, but the log provided is from a consensus node.

I am running the test again at chainhead (obtained with jul-29), but this time without any custom peer connections, only the specified sync/pruning options are applied. I’ll provide an update in a few hours with a new log, but so far it seems that syncing is still failing.

---

<div class="post-metadata">

**Author:** ![nazar-pc](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/nazar-pc/32/6_2.png) [@nazar-pc](https://forum.autonomys.xyz/u/nazar-pc)\
**Post date:** [September 6, 2024, 7:22am UTC](https://forum.autonomys.xyz/t/node-sync-issues-on-sep-03/4409/10 "2024-09-06T07:22:45Z")

</div>

Do you see high CPU usage in the process or not?

---

<div class="post-metadata">

**Author:** ![vexr](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/vexr/32/3498_2.png) [@vexr](https://forum.autonomys.xyz/u/vexr)\
**Post date:** [September 6, 2024, 7:28am UTC](https://forum.autonomys.xyz/t/node-sync-issues-on-sep-03/4409/11 "2024-09-06T07:28:14Z")

</div>

I am seeing high CPU usage. You can see the VM running at \< 3% CPU until I started testing again with sep-03. The system is pegged with this release while trying to sync.

 ![image](https://canada1.discourse-cdn.com/flex011/uploads/subspace/original/2X/2/2d5ac32adaef73d69bcf3ee4aac9e4a19aefe320.png)  
 ![image](https://canada1.discourse-cdn.com/flex011/uploads/subspace/original/2X/e/e5da8908ff01005081c8ef4d2d37ee1034c46136.png)

---

<div class="post-metadata">

**Author:** ![nazar-pc](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/nazar-pc/32/6_2.png) [@nazar-pc](https://forum.autonomys.xyz/u/nazar-pc)\
**Post date:** [September 6, 2024, 7:45am UTC](https://forum.autonomys.xyz/t/node-sync-issues-on-sep-03/4409/12 "2024-09-06T07:45:26Z")

</div>

Interesting, I’ll check it on my end then, large CPU usage indicates it is probably doing something it shouldn’t.

---

<div class="post-metadata">

**Author:** ![vexr](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/vexr/32/3498_2.png) [@vexr](https://forum.autonomys.xyz/u/vexr)\
**Post date:** [September 6, 2024, 9:22am UTC](https://forum.autonomys.xyz/t/node-sync-issues-on-sep-03/4409/13 "2024-09-06T09:22:48Z")

</div>

Here is the latest log. It starts with being able to sync on jul-29 and the restart with sep-03 at timestamp `2024-09-06T06:55:34.248339Z`.

It was never able to fully sync with just over two hours of runtime.

[Log](https://logs.atc.farm/sep-03/node3.log)

---

<div class="post-metadata">

**Author:** ![nazar-pc](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/nazar-pc/32/6_2.png) [@nazar-pc](https://forum.autonomys.xyz/u/nazar-pc)\
**Post date:** [September 9, 2024, 7:00pm UTC](https://forum.autonomys.xyz/t/node-sync-issues-on-sep-03/4409/14 "2024-09-09T19:00:12Z")

</div>

So far can’t reproduce. Had archival node data with height 2430237 blocks (I believe synced with one of the June releases), then synced from DSN using `jul-29`, but it then failed to finalize for some reason, so I switched to `sep-03` and it finished sync shortly afterwards and has no issues staying in sync.

I used `--chain gemini-3h --sync full --blocks-pruning archive-canonical --state-pruning archive-canonical`.
