Repository navigation
FROMLIST: nvme: Add adaptive PCIe link rate switching - #2008
Open
Hariharan Sreedhar (hariharan-sree) wants to merge 6 commits into
Open
Hariharan Sreedhar (hariharan-sree) wants to merge 6 commits into
Hariharan Sreedhar (hariharan-sree) wants to merge 6 commits into
Conversation
added 3 commits
October 9, 2026 17:13
Add support for dynamically adjusting the PCIe link rate of NVMe
devices based on I/O activity. The link is downgraded to a
configurable minimum rate during periods of low activity to save
power, and upgraded to the maximum supported rate as soon as the
workload exceeds a computed threshold, so that peak performance is
maintained under load.
The amount of transferred data is accumulated on the I/O submission
path and aggregated by a per-controller timer. The threshold is
derived from the bandwidth of the minimum and maximum rates such that
the extra time required to transfer the data at the minimum rate does
not exceed the estimated link retraining time.
The feature is currently only supported on Hygon platforms.
Four NVMe devices were tested on the link with this patch. The power
benefit was measured by sampling the voltage and current of the power
monitor every 1 second, each run lasted 3 minutes and the test was
repeated 5 times.
Test run Power Benefit(w)
1 5.3236
2 5.2285
3 5.4176
4 5.4538
5 5.3444
Average 5.3536
The adaptive link rate switching saves about 5.3w of system power
on average.
The feature was tested with Gen1 as the minimum link rate. The workload
runs 10 seconds of random read followed by 0.5 seconds of idle, repeated
100 times.
Avg IOPS (k) Avg BW (GiB/s) Avg clat (us)
With speed switching 2671.2 10.19 378.1
Without speed switching 2681.5 10.23 377.0
The differences are within measurement noise (< 0.5%), which shows
that the adaptive link rate switching does not degrade the baseline
performance and keeps the I/O stable.
In summary, the adaptive link rate switching keeps the I/O performance
essentially unchanged (the differences in IOPS, bandwidth and latency
are all within 0.5% measurement noise) while saving about 5w of
system power, providing an effective way to reduce the platform power
consumption under light I/O without sacrificing throughput or latency.
Signed-off-by: Liao Xuan <liaoxuan@open-hieco.net>
Signed-off-by: Zheng Tan <tanzheng@kylinos.cn>
Link: https://patch.msgid.link/9e99e007bf05a608b1b82da740c2b33935a27175.1790222172.git.liaoxuan@hygon.cn
Signed-off-by: Hariharan Sreedhar <hsreedha@qti.qualcomm.com>
The rate switching decision currently toggles the link rate in every monitoring window based on the I/O of that window alone, which can cause excessive rate switching when the workload fluctuates around the threshold. Add counters that require the threshold to be exceeded for a consecutive number of windows before the rate is upgraded, and to stay below the threshold for a consecutive number of windows before it is downgraded. The upgrade happens immediately (threshold 0 by default) so that latency is not sacrificed, while the downgrade requires 10 consecutive windows. Monitoring is stopped after 10 seconds of inactivity and is re-armed by the next I/O. Signed-off-by: Liao Xuan <liaoxuan@open-hieco.net> Link: https://patch.msgid.link/de23e8b9699b5a7a77e7ee215e8e4c9f530ce5da.1790222172.git.liaoxuan@hygon.cn Signed-off-by: Hariharan Sreedhar <hsreedha@qti.qualcomm.com>
Add a speed attribute group under /sys/class/nvme/nvmeX/ that exposes the adaptive link rate switching parameters for runtime tuning: enable - toggle the feature on/off (0/1) monitor_interval - I/O monitoring window in milliseconds (>= 100) min_speed - minimum PCIe generation to downgrade to (1-5) up_threshold - windows above the threshold needed to upgrade down_threshold - windows below the threshold needed to downgrade The attributes are read/write and are removed when the controller is torn down. Signed-off-by: Liao Xuan <liaoxuan@open-hieco.net> Link: https://patch.msgid.link/be224cd8385a5c28a20db65d635000a7be4da91a.1790222172.git.liaoxuan@hygon.cn Signed-off-by: Hariharan Sreedhar <hsreedha@qti.qualcomm.com>
qcomlnxci
requested review from
a team,
krishnachaitanya-linux and
Matthew Leung (meleung)
and removed request for
a team
October 9, 2026 11:57
… around link retraining PCIe host bridge controllers may need their operating point raised before retraining to a higher link speed so that hardware resources (e.g., RPMh votes on Qualcomm platforms) are available at the requested data rate. After retraining, the operating point must be updated to reflect the actual negotiated speed. Add pcie_set_opp() to look up an OPP on the host bridge parent device using a key of (per-lane frequency in kHz, LNKCTL2 Target Link Speed level). Keying by generation rather than total bandwidth lets OPP tables remain width-independent. In pcie_set_target_speed(), call pcie_set_opp() before retraining only when upscaling (speed_req > cur_bus_speed), since only raising the operating point requires pre-staging hardware. After retraining, call pcie_set_opp() unconditionally with the actual cur_bus_speed to settle the votes. Both calls are skipped for downstream ports of PCIe switches, as those are outside the host controller's scope. Some controllers also require ASPM to be disabled around link retraining. Add a disable_aspm_for_retrain flag to pci_host_bridge; when set, pcie_set_target_speed() saves the child device's ASPM state, disables all ASPM link states before retraining, and restores them afterward. Link: https://patch.msgid.link/20260819-bwscale-v5-1-6dea79786b37@oss.qualcomm.com Signed-off-by: Krishna Chaitanya Chundru <krishna.chundru@oss.qualcomm.com> Signed-off-by: Hariharan Sreedhar <hsreedha@qti.qualcomm.com>
Set the disable_aspm_for_retrain flag in qcom_pcie_host_init() to ensure ASPM is disabled during link retraining operations. This prevents potential issues with link state transitions during bandwidth scaling on Qualcomm PCIe controllers. Signed-off-by: Krishna Chaitanya Chundru <krishnac@codeaurora.org> Link: https://patch.msgid.link/20260819-bwscale-v5-6-6dea79786b37@oss.qualcomm.com Signed-off-by: Hariharan Sreedhar <hsreedha@qti.qualcomm.com>
The Qualcomm PCIe controller hardwires the Link Bandwidth Notification Capability (LBNC) bit to 0, even though the Root Port supports retraining the link to the speeds advertised in LNKCAP2. The PCIe port bandwidth controller relies on this capability to determine whether link bandwidth control is supported. With LBNC cleared, the bandwidth controller is not registered and link speed cannot be managed through the associated thermal cooling device. Override the read-only LNKCAP register through the DBI RO write interface and set PCI_EXP_LNKCAP_LBNC during controller initialization. This allows the PCIe bandwidth controller to bind to the Root Port and enables link speed scaling through the thermal framework. Signed-off-by: Krishna Chaitanya Chundru <krishna.chundru@oss.qualcomm.com> Signed-off-by: Hariharan Sreedhar <hsreedha@qti.qualcomm.com>
Hariharan Sreedhar (hariharan-sree)
force-pushed
the
nvme_dynamic_BW_scale
branch
from
October 9, 2026 12:40
9fbac67 to
579ae3e
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Add support for dynamically adjusting the PCIe link rate of NVMe devices based on I/O activity. The link is downgraded to a
configurable minimum rate during periods of low activity to save power, and upgraded to the maximum supported rate as soon as the workload exceeds a computed threshold, so that peak performance is maintained under load.
The feature is currently only supported on Hygon platforms.
Link : https://patch.msgid.link/9e99e007bf05a608b1b82da740c2b33935a27175.1790222172.git.liaoxuan@hygon.cn