Movatterモバイル変換

kouvel added this to the6.0.0 milestone

kouvel requested review fromdavidfowl,janvorli andmangod9

May 30, 2021 01:37

kouvel self-assigned this

Copy link

kouvel commentedMay 30, 2021•
edited
Loading

Checked perf on thread pool overhead tests on x64 and arm64, no significant difference. Checked perf on ASP.NET platform benchmarks on x64, no significant difference.
Verified config vars are working as expected
Verified throttling rate of thread injection in low-memory situations with Windows job objects and Linux docker containers
Checked some cases involving interaction with starvation and hill climbing heuristics, verified behavior is appropriate. Solution is not quite ideal until we also fix the starvation heuristic, but I tried to make sure that the likelihood of a new issue is low.
Checked the relevant cases fromhttps://github.com/davidfowl/AspNetCoreDiagnosticScenarios and verified that the thread pool is more responsive to compensate for the sync-over-async blocking work
The defaults for config vars are resulting from a reasonable guess arising from brief prior discussions on the topic, there are good reasons for the limits, but we are also not trying to make every real-world really-bad scenario involving sync-over-async work really-well by default. It's a realistic expectation that the really bad cases would involve some configuration. An expectation in those cases where sync-over-async is the only type of blocking happening on thread pool worker threads, is that the new config vars would work better than setting a high min worker thread count, because this solution uses cooperative blocking and can adjust active thread counts up and down appropriately. Setting a high min thread count on the other hand is not a great workaround for that problem because it causes that many threads to always be active, and that's not ideal.

kouvel requested a review frommarek-safar as acode owner

May 30, 2021 02:14

Copy link

Member

davidfowl commentedMay 30, 2021•
edited
Loading

What do you think about doing this in monitor.wait as well?

Copy link

ContributorAuthor

kouvel commentedMay 30, 2021

What do you think about doing this in monitor.wait as well?

In my opinion, a monitor is too basic of a synchronization primitive to assume that waiting on one would always deserve compensating for. For instance, it's often not beneficial or preferable to add threads to compensate for threads blocking on waiting to acquire a lock.

benaadams reviewed

UpdateThreadCounts usage based on a changedotnet/diagnostics#2324

src/libraries/System.Private.CoreLib/src/System/Threading/ThreadPoolWorkQueue.cs OutdatedShow resolvedHide resolved

This was referencedMay 31, 2021

Merged

Add a thread adjustment reasonmicrosoft/perfview#1439

Merged

stephentoub reviewed

src/libraries/System.Private.CoreLib/src/System/Threading/ThreadPoolWorkQueue.cs OutdatedShow resolvedHide resolved

stephentoub reviewed

src/libraries/System.Private.CoreLib/src/System/Threading/ThreadPoolWorkQueue.cs OutdatedShow resolvedHide resolved

stephentoub reviewed

src/libraries/System.Private.CoreLib/src/System/Threading/PortableThreadPool.Blocking.cs OutdatedShow resolvedHide resolved

stephentoub reviewed

src/libraries/System.Private.CoreLib/src/System/Threading/PortableThreadPool.Blocking.cs OutdatedShow resolvedHide resolved

stephentoub reviewed

src/libraries/System.Private.CoreLib/src/System/Threading/PortableThreadPool.Blocking.cs OutdatedShow resolvedHide resolved

stephentoub reviewed

src/libraries/System.Private.CoreLib/src/System/Threading/PortableThreadPool.Blocking.cs OutdatedShow resolvedHide resolved

stephentoub reviewed

src/libraries/System.Private.CoreLib/src/System/Threading/PortableThreadPool.Blocking.cs OutdatedShow resolvedHide resolved

mangod9 reviewed

src/libraries/System.Private.CoreLib/src/System/Threading/PortableThreadPool.WorkerThread.cs OutdatedShow resolvedHide resolved

mangod9 reviewed

src/libraries/System.Private.CoreLib/src/System/Threading/PortableThreadPool.ThreadCounts.cs OutdatedShow resolvedHide resolved

stephentoub reviewed

src/libraries/System.Private.CoreLib/src/System/Threading/PortableThreadPool.Blocking.cs Outdated

Copy link

Member

stephentoubJun 1, 2021

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others.Learn more.

There's a lot of policy in here, and a lot of knobs to go with it. Do you have a sense for how all of this is going to behave in the real-world, and if/how someone would utilize these knobs effectively? How did you arrive at this specific set and also the defaults employed?

Copy link

ContributorAuthor

kouvelJun 2, 2021•
edited
Loading

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others.Learn more.

Some of the criteria used:

Have a good replacement for setting theMinThreads as a workaround
- This would now be to setThreadsToAddWithoutDelay to the equivalent and setMaxThreadsToAddBeforeFallback to something higher to give some buffer for spikes that may need more threads
- MaxThreadsToAddBeforeFallback could also be set to a large value to effectively unlimit the heuristic
Use progressive delays to avoid creating too many threads too quickly
- Without that, it would be conceivable thatMaxThreadsToAddBeforeFallback threads would be created in short order to respond to even a short sync-over-async IO before the IO even completes (if there are that many work items that would block on the async work)
- The delay also helps to create a high watermark of how many threads were necessary last time to unblock, so that when there's a limit to how many work items would block, it would quickly release existing waiting threads to let other work be done meanwhile
- The larger the number of threads that get unblocked all at once, the higher the latency of their processing would be after unblocking. There's probably not a good solution to this.
- Ideally it would not require as many threads for an async IO completion to unblock waiting threads, sort of like on Windows where a separate pool of threads handles IO completions, needs some experimentation
- It's not always clear that adding more threads would help, more so in starvation-type cases
No one set of defaults will work well for all cases, use conservative defaults to start with
- The current defaults are much more agressive than before
- TheMaxThreadsToAddBeforeFallback_ProcCountFactor of10 came from a prior discussion where we felt that adding 10x the proc count relatively quickly may not be too bad
- The defaults can be made more aggressive easily, but it would be difficult to make the defaults more conservative since apps that work well with the defaults may not after that without configuration
- The really bad cases where many 100s or even 1000s of threads need to be created will likely need to configure for the app's situation based on expected workload and how bad it can get, in order to work around the issue
Make things sufficiently configurable
- It would have been nice to make configurable the delay threshold for detecting starvation and the delay used to add threads during continuous starvation. Now, for sync-over-async the delay and rate of progression in delays can be adjusted.
- Similarly to hill climbing config values, the config values don't have to be used but it can be helpful to enable the freedom to configure them
- I expect I would suggest most users running into bad blocking due to sync-over-async to configureThreadsToAddWithoutDelay andMaxThreadsToAddBeforeFallback, and perhapsMaxDelayUntilFallbackMs to control the thread injection latency for spikes
I intend to use the same config values (maybe with a couple of others) for improving the starvation heuristic similarly in the future
- Starvation is a bit different and may need a few quirks, but hopefully we can use something similar

Copy link

ContributorAuthor

kouvelJun 4, 2021

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others.Learn more.

Decided to removeMaxThreadsToAddBeforeFallback and renamedMaxDelayBeforeFallbackMs toMaxDelayMs in the latest commit. The max threads limit before falling back to starvation seems unnecessary, it would be unlimited for starvation anyway. Now the only time it would fall back to starvation is in low-memory situations.

janvorli reviewed

Host tests often fail on text file busy issues#53587

src/libraries/System.Private.CoreLib/src/System/Threading/PortableThreadPool.ThreadCounts.cs OutdatedShow resolvedHide resolved

src/libraries/System.Private.CoreLib/src/System/Threading/PortableThreadPool.cs OutdatedShow resolvedHide resolved

runfoappbot mentioned this pull request

Jun 2, 2021

Closed

Koundinya Veluri added4 commits

June 2, 2021 10:29

Improve the rate of thread injection for blocking due to sync-over-async

55af247

Fixes#52558

Fix browser build

726dc9e

Remove an invalid assertion

96cc50e

Fix bad merge

f3e8a31

Koundinya Veluri added5 commits

June 2, 2021 10:29

Update event enum with new thread adjustment reason

d3a977b

Log all of the transitions for now (will change later when needed), a…

c2cb717

…nd fix throughput numbers sent in events

Add a test for some coverage

20526d6

Fix test

984888f

Address feedback, add an assertion

56bd4d9

Copy link

ContributorAuthor

kouvel commentedJun 2, 2021

Rebased to fix conflicts

Koundinya Veluri added3 commits

June 2, 2021 11:03

Fix browser build

8fc1f8a

Remove max threads config for fallback, rename max delay config

1fe2067

Actually rename config

e7b9508

Copy link

ContributorAuthor

kouvel commentedJun 7, 2021

I believe I have addressed the feedback shared so far, any other feedback?

mangod9 approved these changes

Jun 8, 2021