# Is this NUMA test result real?

**URL:** <https://forum.autonomys.xyz/t/is-this-numa-test-result-real/2323>\
**Category:** Support\
**Created:** [January 3, 2024, 7:24am UTC](https://forum.autonomys.xyz/t/is-this-numa-test-result-real/2323 "2024-01-03T07:24:33Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![z\_W](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/z_w/32/1135_2.png) [@z\_W](https://forum.autonomys.xyz/u/z_W)\
**Post date:** [January 3, 2024, 7:24am UTC](https://forum.autonomys.xyz/t/is-this-numa-test-result-real/2323/1 "2024-01-03T07:24:33Z")

</div>

1. CPU AMD Epyc 7302 x2, RAM 16G DDR4 x16 _(meaning motherboard has two Epyc 7302 processors installed and 16 memory modules, 16G each)_
2. 5m10s per sector _(one sector is encoded at a time by default)_
3. 6m0s per sector, 4 sectors at a time _(meaning number of downloaded and encoded sectors was manually increased)_
4. 4m30s per sector, 8 sectors at a time _(meaning 8 NUMA nodes)_
5. 2m90s per sector, 8 sectors at a time _(meaning 8 NUMA no_

---

<div class="post-metadata">

**Author:** ![nazar-pc](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/nazar-pc/32/6_2.png) [@nazar-pc](https://forum.autonomys.xyz/u/nazar-pc)\
**Post date:** [January 3, 2024, 11:49am UTC](https://forum.autonomys.xyz/t/is-this-numa-test-result-real/2323/2 "2024-01-03T11:49:59Z")

</div>

Looks plausible, though impact of NUMA-aware memory allocator seems huge, especially considering that 7002 Epyc processors use I/O die and I don’t think there should be significant difference between accessing any of the memory channels, though you do have two physical sockets and maybe crossing from one socket to another is very costly on 7002 Epyc platform.

If after repeated tests this is confirmed, we might make NUMA-aware memory allocator the default because negative impact on other platforms is limited and benefit here is massive.

---

<div class="post-metadata">

**Author:** ![z\_W](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/z_w/32/1135_2.png) [@z\_W](https://forum.autonomys.xyz/u/z_W)\
**Post date:** [January 3, 2024, 12:11pm UTC](https://forum.autonomys.xyz/t/is-this-numa-test-result-real/2323/3 "2024-01-03T12:11:30Z")

</div>

The optimization in the new version is still not as fast as running multiple instances of the software.

---

<div class="post-metadata">

**Author:** ![nazar-pc](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/nazar-pc/32/6_2.png) [@nazar-pc](https://forum.autonomys.xyz/u/nazar-pc)\
**Post date:** [January 3, 2024, 12:16pm UTC](https://forum.autonomys.xyz/t/is-this-numa-test-result-real/2323/4 "2024-01-03T12:16:46Z")

</div>

You numbers say the opposite though 🤔

---

<div class="post-metadata">

**Author:** ![z\_W](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/z_w/32/1135_2.png) [@z\_W](https://forum.autonomys.xyz/u/z_W)\
**Post date:** [January 3, 2024, 12:28pm UTC](https://forum.autonomys.xyz/t/is-this-numa-test-result-real/2323/5 "2024-01-03T12:28:17Z")

</div>

This is the test data you gave.

> [@NUMA support is coming](https://forum.autonomys.xyz/t/numa-support-is-coming/2299):
>
> For a long time farmers were saying that plotting is slow on large CPUs, now it is time to change that! I’ve been hacking on NUMA support that should make things much better and need folks to test and provide feedback to confirm it is actually a positive change. Please read this post to the very end before replying! What is changing There are several behaviors on the farmer that will be different. Global thread pools Previously plotting/replotting thread pools were created for each farm sepa…

---

<div class="post-metadata">

**Author:** ![nazar-pc](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/nazar-pc/32/6_2.png) [@nazar-pc](https://forum.autonomys.xyz/u/nazar-pc)\
**Post date:** [January 3, 2024, 12:30pm UTC](https://forum.autonomys.xyz/t/is-this-numa-test-result-real/2323/6 "2024-01-03T12:30:47Z")

</div>

> [@z\_W](#):
>
> This is the test data you gave.

I mean in your first message version of the farmer with NUMA support is faster than version without NUMA support (even when configured to plot 4 sectors at a time). Why are you saying it is not as fast as running multiple instances?

---

<div class="post-metadata">

**Author:** ![z\_W](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/z_w/32/1135_2.png) [@z\_W](https://forum.autonomys.xyz/u/z_W)\
**Post date:** [January 3, 2024, 12:34pm UTC](https://forum.autonomys.xyz/t/is-this-numa-test-result-real/2323/7 "2024-01-03T12:34:55Z")

</div>

I have a server with only two NUMA nodes. Running the test version, the speed is 5m-6m\*2, but I can only open one software. When I open two, the speed drops to over 10 minutes.

Without using the NUMA version, I can open four software, each running stably at 7m\*1.

EPYC7302\*2 is your CPU

---

<div class="post-metadata">

**Author:** ![nazar-pc](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/nazar-pc/32/6_2.png) [@nazar-pc](https://forum.autonomys.xyz/u/nazar-pc)\
**Post date:** [January 3, 2024, 12:40pm UTC](https://forum.autonomys.xyz/t/is-this-numa-test-result-real/2323/8 "2024-01-03T12:40:24Z")

</div>

Hm… the whole point of the new version is to utilize CPU fully, you shouldn’t need more than one instance because it’ll be less efficient, which is exactly what you see. Running multiple instances was a workaround for not supporting NUMA that is no longer necessary.

---

<div class="post-metadata">

**Author:** ![nazar-pc](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/nazar-pc/32/6_2.png) [@nazar-pc](https://forum.autonomys.xyz/u/nazar-pc)\
**Post date:** [January 3, 2024, 12:46pm UTC](https://forum.autonomys.xyz/t/is-this-numa-test-result-real/2323/9 "2024-01-03T12:46:33Z")

</div>

Ah, sorry for confusion. Those were just examples, they are made up numbers and just provided for illustration purposes to show how to submit test results.

---

<div class="post-metadata">

**Author:** ![z\_W](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/z_w/32/1135_2.png) [@z\_W](https://forum.autonomys.xyz/u/z_W)\
**Post date:** [January 3, 2024, 1:10pm UTC](https://forum.autonomys.xyz/t/is-this-numa-test-result-real/2323/10 "2024-01-03T13:10:40Z")

</div>

My CPU has many cores, but there are only two NUMA nodes. I expect the ideal speed for my CPU to be 7m-8m_4. However, the actual speed is 5m-6m_2.

---

<div class="post-metadata">

**Author:** ![nazar-pc](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/nazar-pc/32/6_2.png) [@nazar-pc](https://forum.autonomys.xyz/u/nazar-pc)\
**Post date:** [January 3, 2024, 1:18pm UTC](https://forum.autonomys.xyz/t/is-this-numa-test-result-real/2323/11 "2024-01-03T13:18:47Z")

</div>

> [@z\_W](#):
>
> I expect the ideal speed for my CPU to be 7m-8m_4_

Why?

> [@z\_W](#):
>
> _However, the actual speed is 5m-6m_2

Isn’t this not a good thing?

---

<div class="post-metadata">

**Author:** ![z\_W](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/z_w/32/1135_2.png) [@z\_W](https://forum.autonomys.xyz/u/z_W)\
**Post date:** [January 3, 2024, 1:39pm UTC](https://forum.autonomys.xyz/t/is-this-numa-test-result-real/2323/12 "2024-01-03T13:39:28Z")

</div>

Assuming I have 4 SSDs, the speed when running multiple instances of the software is 1SSD 7m-8m x1 x4. Using the new version and opening only one instance of the software, the speed is 4SSD 5m-6m x2 x1.

---

<div class="post-metadata">

**Author:** ![nazar-pc](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/nazar-pc/32/6_2.png) [@nazar-pc](https://forum.autonomys.xyz/u/nazar-pc)\
**Post date:** [January 3, 2024, 2:02pm UTC](https://forum.autonomys.xyz/t/is-this-numa-test-result-real/2323/13 "2024-01-03T14:02:25Z")

</div>

Did you customize any CLI options related to thread pool size or number of encoded sectors in new version? They will interfere with the intended behavior.

Overall it is possible that on some CPUs it will still be non-ideal in some configurations, for example in case of your 8272CL it is simply not the most optimal CPU due to just 2 NUMA nodes and such a massive number of cores in each that many algorithms will not take full advantage of it.

You should still be able to benefit by running two instances instead of 4. In worst case you’ll just run 4 instances like before. As long as performance doesn’t regress I think it is a win because NUMA support is clearly better than previous default.

BTW, with new version threads are pinned to cores, so if you specify encoding concurrency to `4`, you should get very good CPU utilization while also avoiding crossing NUMA nodes with just one farmer.

---

<div class="post-metadata">

**Author:** ![z\_W](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/z_w/32/1135_2.png) [@z\_W](https://forum.autonomys.xyz/u/z_W)\
**Post date:** [January 3, 2024, 2:35pm UTC](https://forum.autonomys.xyz/t/is-this-numa-test-result-real/2323/14 "2024-01-03T14:35:17Z")

</div>

> [@nazar-pc](#):
>
> Did you customize any CLI options related to thread pool size or number of encoded sectors in new version? They will interfere with the intended behavior.

When I start a software process, I use the default parameters.

> [@nazar-pc](#):
>
> You should still be able to benefit by running two instances instead of 4. In worst case you’ll just run 4 instances like before. As long as performance doesn’t regress I think it is a win because NUMA support is clearly better than previous default.
> 
> BTW, with new version threads are pinned to cores, so if you specify encoding concurrency to `4`, you should get very good CPU utilization while also avoiding crossing NUMA nodes with just one farmer.

I am trying to start two software processes. Can I start two processes with these parameters?

```auto
--sector-downloading-concurrency 4 
--sector-encoding-concurrency 4
--farming-thread-pool-size 10
--plotting-thread-pool-size 16
--replotting-thread-pool-size 8

```

so?

---

<div class="post-metadata">

**Author:** ![nazar-pc](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/nazar-pc/32/6_2.png) [@nazar-pc](https://forum.autonomys.xyz/u/nazar-pc)\
**Post date:** [January 3, 2024, 2:59pm UTC](https://forum.autonomys.xyz/t/is-this-numa-test-result-real/2323/15 "2024-01-03T14:59:22Z")

</div>

If you want to have 4 farms plotted at the same time with one instance of the farmer, I would recommend to specify a single option:

```auto
--sector-encoding-concurrency 4

```

Farmer should be able to calculate all other options automatically in an optimal way.

This will result in half of each CPU being dedicated to plotting of a single sector, replotting will be configured to 1/4 of CPU core and downloading concurrency will be set to optimal value of 5. If you want overlap between sectors for plotting process you might also add `--plotting-thread-pool-size 52` and each NUMA node will be processing 2 farms at the same time, but there will still be no NUMA node crossing.

I’m fairly certain it will be more efficient than running multiple farmer instances, especially if you’re not pinning them to NUMA nodes.

---

<div class="post-metadata">

**Author:** ![z\_W](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/z_w/32/1135_2.png) [@z\_W](https://forum.autonomys.xyz/u/z_W)\
**Post date:** [January 3, 2024, 3:04pm UTC](https://forum.autonomys.xyz/t/is-this-numa-test-result-real/2323/16 "2024-01-03T15:04:29Z")

</div>

> [@nazar-pc](#):
>
> –plotting-thread-pool-size 52

I’ll try the parameters you recommended

---

<div class="post-metadata">

**Author:** ![z\_W](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/z_w/32/1135_2.png) [@z\_W](https://forum.autonomys.xyz/u/z_W)\
**Post date:** [January 3, 2024, 3:49pm UTC](https://forum.autonomys.xyz/t/is-this-numa-test-result-real/2323/17 "2024-01-03T15:49:01Z")

</div>

![image](https://canada1.discourse-cdn.com/flex011/uploads/subspace/original/2X/2/2086d905a0ec7dc0bd27ec56f93f72d1a47c242f.jpeg)

---

<div class="post-metadata">

**Author:** ![nazar-pc](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/nazar-pc/32/6_2.png) [@nazar-pc](https://forum.autonomys.xyz/u/nazar-pc)\
**Post date:** [January 3, 2024, 4:12pm UTC](https://forum.autonomys.xyz/t/is-this-numa-test-result-real/2323/18 "2024-01-03T16:12:24Z")

</div>

So ~4m45s per sector, not very fast for such system. You can set number of plotting threads to 104 to achieve the same result as running multiple separate farmers. This is up to you to experiment and share the findings.

---

<div class="post-metadata">

**Author:** ![z\_W](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/z_w/32/1135_2.png) [@z\_W](https://forum.autonomys.xyz/u/z_W)\
**Post date:** [January 4, 2024, 4:18am UTC](https://forum.autonomys.xyz/t/is-this-numa-test-result-real/2323/19 "2024-01-04T04:18:14Z")

</div>

> [@nazar-pc](#):
>
> plotting threads to 104

more slowly,

–farming-thread-pool-size 10  
–plotting-thread-pool-size 16  
–replotting-thread-pool-size 8

I usually use this parameter to start four software processes

---

<div class="post-metadata">

**Author:** ![z\_W](https://yyz1.discourse-cdn.com/flex011/user_avatar/forum.autonomys.xyz/z_w/32/1135_2.png) [@z\_W](https://forum.autonomys.xyz/u/z_W)\
**Post date:** [January 7, 2024, 3:27pm UTC](https://forum.autonomys.xyz/t/is-this-numa-test-result-real/2323/20 "2024-01-07T15:27:24Z")

</div>

```auto
2024-01-07T15:25:39.261781Z INFO subspace_farmer::commands::farm: NUMA system detected numa_nodes=2
2024-01-07T15:25:39.261794Z WARN subspace_farmer::commands::farm: Too few disk farms, CPU will not be utilized fully during plotting, same number of farms as NUMA nodes or more is recommended numa_nodes=2 farms_count=4

```

why is that?

[Next page](https://forum.autonomys.xyz/t/is-this-numa-test-result-real/2323.md?page=2)
