From patchwork Mon Jun 5 14:14:41 2023 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 8bit X-Patchwork-Submitter: tip-bot2 for Thomas Gleixner X-Patchwork-Id: 103306 Return-Path: Delivered-To: ouuuleilei@gmail.com Received: by 2002:a59:994d:0:b0:3d9:f83d:47d9 with SMTP id k13csp2735266vqr; Mon, 5 Jun 2023 07:44:35 -0700 (PDT) X-Google-Smtp-Source: ACHHUZ7pjZPzipf7+wQ7QZYAi8tbmoQomSkk4WgbrQKz6dqOPEckrAcpR07TgwQqcgW9alR9MtdP X-Received: by 2002:a17:902:c20c:b0:1b0:348:2498 with SMTP id 12-20020a170902c20c00b001b003482498mr7728520pll.2.1685976274916; Mon, 05 Jun 2023 07:44:34 -0700 (PDT) ARC-Seal: i=1; a=rsa-sha256; t=1685976274; cv=none; d=google.com; s=arc-20160816; b=mnHLffgONBkriHLiZXe3aDfBRPtnP9SJR2kq5gMNSqGcz2CU2KWbWTVEVmKwmjizfd J72khoF+SMwjK+Fzy98R+yc7ZOMOXlVqaL+XiqvKQrSYKSqExtqp12HIKeESvaJdWuGM YCIkTbRmTaEb0N4qq6N3NbGYS1qw1PJJ+6PWfnhPMI48SO/2nUGNT5mlRFThDtoFD2Ub zsHRHjct0QheVxEvYz1UCfb1dHjOMFf7ja6dunWUXXflwD1PN8oYU1ai475dcn9CrNGa J4OBuu2pMAnuiXcNrSlx5tj+muOWroOUa++YCO4cQZVd0+dPhicmY+X3tSLzifGUiZ+G hRFA== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=list-id:precedence:content-transfer-encoding:robot-unsubscribe :robot-id:message-id:mime-version:references:in-reply-to:cc:subject :to:reply-to:sender:from:dkim-signature:dkim-signature:date; bh=KDsnAbFl/koTE0Dl0BN3mQjS+9ciwBGJ3f8Zlh4UZLQ=; b=06hPqVbv/g05K1LBO+flDpv7Ij+7lxKkgDMEZGh69V3GJlDhvJzH3iBKsMqxlSvMHp Y5ksKvKMnxT3jU4xkWKXHFs4BOF1zDJo8YCnufN+JGfUkCM3SoaR6AYLQw8uhso1/nfp phH8HjetnsNhXdC2m3FOOpPR/I+r6GWOcAUCSe3ZbMNQTu9kkEfyPQ8UpwDEO/EB3CyQ 6bBvuifEUXgf5y0uzxZYxiTSDBp+wqSQEOzDKPiP51RHwOITqovKJQdD3welqcmpKrNH tkTpPSn+XAdtnRdhS0lC5P+kLUok/kCpHANg5UC3fvU4ErPcFiBGoBP8Fx+NOKVSImoe o0TA== ARC-Authentication-Results: i=1; mx.google.com; dkim=pass header.i=@linutronix.de header.s=2020 header.b=qK5R6FrY; dkim=neutral (no key) header.i=@linutronix.de header.s=2020e header.b="OMuRKr3/"; spf=pass (google.com: domain of linux-kernel-owner@vger.kernel.org designates 2620:137:e000::1:20 as permitted sender) smtp.mailfrom=linux-kernel-owner@vger.kernel.org; dmarc=pass (p=NONE sp=QUARANTINE dis=NONE) header.from=linutronix.de Received: from out1.vger.email (out1.vger.email. [2620:137:e000::1:20]) by mx.google.com with ESMTP id ja19-20020a170902efd300b001ae35b8d593si5378582plb.264.2023.06.05.07.44.19; Mon, 05 Jun 2023 07:44:34 -0700 (PDT) Received-SPF: pass (google.com: domain of linux-kernel-owner@vger.kernel.org designates 2620:137:e000::1:20 as permitted sender) client-ip=2620:137:e000::1:20; Authentication-Results: mx.google.com; dkim=pass header.i=@linutronix.de header.s=2020 header.b=qK5R6FrY; dkim=neutral (no key) header.i=@linutronix.de header.s=2020e header.b="OMuRKr3/"; spf=pass (google.com: domain of linux-kernel-owner@vger.kernel.org designates 2620:137:e000::1:20 as permitted sender) smtp.mailfrom=linux-kernel-owner@vger.kernel.org; dmarc=pass (p=NONE sp=QUARANTINE dis=NONE) header.from=linutronix.de Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S233033AbjFEOUr (ORCPT + 99 others); Mon, 5 Jun 2023 10:20:47 -0400 Received: from lindbergh.monkeyblade.net ([23.128.96.19]:40484 "EHLO lindbergh.monkeyblade.net" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S234911AbjFEOPJ (ORCPT ); Mon, 5 Jun 2023 10:15:09 -0400 Received: from galois.linutronix.de (Galois.linutronix.de [193.142.43.55]) by lindbergh.monkeyblade.net (Postfix) with ESMTPS id AC51C10F9; Mon, 5 Jun 2023 07:14:43 -0700 (PDT) Date: Mon, 05 Jun 2023 14:14:41 -0000 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linutronix.de; s=2020; t=1685974482; h=from:from:sender:sender:reply-to:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=KDsnAbFl/koTE0Dl0BN3mQjS+9ciwBGJ3f8Zlh4UZLQ=; b=qK5R6FrY9BZwV5VW7Zo4ZQD2/2RSZ3xMrI3pvRI/CL3mbJ8jQ2yu/iQhz9J5xsDP06D6aN n67lXeM7aHSqGV3tCc15bxAFSSkMdiovK5Tj6pIGRxbifi7uVm02YWfqslOQNZe/qdKYom Pn3YSrH4/UBUQzCr3hUgk35HxCIX+w7AgimXdt+XEAr5ibvrcKpjwsuV3uimoQLLHPan4V GPgAUEhqwbnp8Vjd2gkJRyST6qN19k+WW6+5mv78/oEOalEMrfkypqqgq+WGrWxfzHiGUJ p/2C29Gy/ILNHXreYvMuSmGXR02064sa0Hc/aB1ebe4EG029k/RrlnpYbCAXSg== DKIM-Signature: v=1; a=ed25519-sha256; c=relaxed/relaxed; d=linutronix.de; s=2020e; t=1685974482; h=from:from:sender:sender:reply-to:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=KDsnAbFl/koTE0Dl0BN3mQjS+9ciwBGJ3f8Zlh4UZLQ=; b=OMuRKr3/W8bytphTE1gaGXa63Er0Z+TuyB4b+e2Ifh0fDyNhdJwbjRZC+ZcrRJrnD6GkYn Zocc6eAsmBJyQuAQ== From: "tip-bot2 for Muralidhara M K" Sender: tip-bot2@linutronix.de Reply-to: linux-kernel@vger.kernel.org To: linux-tip-commits@vger.kernel.org Subject: [tip: ras/core] EDAC/amd64: Document heterogeneous system enumeration Cc: Muralidhara M K , Naveen Krishna Chatradhi , Yazen Ghannam , "Borislav Petkov (AMD)" , x86@kernel.org, linux-kernel@vger.kernel.org In-Reply-To: <20230515113537.1052146-4-muralimk@amd.com> References: <20230515113537.1052146-4-muralimk@amd.com> MIME-Version: 1.0 Message-ID: <168597448189.404.1787278187429193729.tip-bot2@tip-bot2> Robot-ID: Robot-Unsubscribe: Contact to get blacklisted from these emails X-Spam-Status: No, score=-4.4 required=5.0 tests=BAYES_00,DKIM_SIGNED, DKIM_VALID,DKIM_VALID_AU,DKIM_VALID_EF,RCVD_IN_DNSWL_MED,SPF_HELO_NONE, SPF_PASS,T_SCC_BODY_TEXT_LINE autolearn=ham autolearn_force=no version=3.4.6 X-Spam-Checker-Version: SpamAssassin 3.4.6 (2021-04-09) on lindbergh.monkeyblade.net Precedence: bulk List-ID: X-Mailing-List: linux-kernel@vger.kernel.org X-getmail-retrieved-from-mailbox: =?utf-8?q?INBOX?= X-GMAIL-THRID: =?utf-8?q?1765960973915581078?= X-GMAIL-MSGID: =?utf-8?q?1767874258311843830?= The following commit has been merged into the ras/core branch of tip: Commit-ID: 4f3fa571a48feb56e7ed1978a27983b89dd2107a Gitweb: https://git.kernel.org/tip/4f3fa571a48feb56e7ed1978a27983b89dd2107a Author: Muralidhara M K AuthorDate: Mon, 15 May 2023 11:35:35 Committer: Borislav Petkov (AMD) CommitterDate: Mon, 05 Jun 2023 12:27:15 +02:00 EDAC/amd64: Document heterogeneous system enumeration Document High Bandwidth Memory (HBM) and AMD heterogeneous system topology and enumeration. [ bp: Simplify and de-marketize, unify, massage. ] Signed-off-by: Muralidhara M K Co-developed-by: Naveen Krishna Chatradhi Signed-off-by: Naveen Krishna Chatradhi Signed-off-by: Yazen Ghannam Signed-off-by: Borislav Petkov (AMD) Link: https://lore.kernel.org/r/20230515113537.1052146-4-muralimk@amd.com --- Documentation/driver-api/edac.rst | 120 +++++++++++++++++++++++++++++- 1 file changed, 120 insertions(+) diff --git a/Documentation/driver-api/edac.rst b/Documentation/driver-api/edac.rst index b8c742a..f4f044b 100644 --- a/Documentation/driver-api/edac.rst +++ b/Documentation/driver-api/edac.rst @@ -106,6 +106,16 @@ will occupy those chip-select rows. This term is avoided because it is unclear when needing to distinguish between chip-select rows and socket sets. +* High Bandwidth Memory (HBM) + +HBM is a new memory type with low power consumption and ultra-wide +communication lanes. It uses vertically stacked memory chips (DRAM dies) +interconnected by microscopic wires called "through-silicon vias," or +TSVs. + +Several stacks of HBM chips connect to the CPU or GPU through an ultra-fast +interconnect called the "interposer". Therefore, HBM's characteristics +are nearly indistinguishable from on-chip integrated RAM. Memory Controllers ------------------ @@ -176,3 +186,113 @@ nodes:: the L1 and L2 directories would be "edac_device_block's" .. kernel-doc:: drivers/edac/edac_device.h + + +Heterogeneous system support +---------------------------- + +An AMD heterogeneous system is built by connecting the data fabrics of +both CPUs and GPUs via custom xGMI links. Thus, the data fabric on the +GPU nodes can be accessed the same way as the data fabric on CPU nodes. + +The MI200 accelerators are data center GPUs. They have 2 data fabrics, +and each GPU data fabric contains four Unified Memory Controllers (UMC). +Each UMC contains eight channels. Each UMC channel controls one 128-bit +HBM2e (2GB) channel (equivalent to 8 X 2GB ranks). This creates a total +of 4096-bits of DRAM data bus. + +While the UMC is interfacing a 16GB (8high X 2GB DRAM) HBM stack, each UMC +channel is interfacing 2GB of DRAM (represented as rank). + +Memory controllers on AMD GPU nodes can be represented in EDAC thusly: + + GPU DF / GPU Node -> EDAC MC + GPU UMC -> EDAC CSROW + GPU UMC channel -> EDAC CHANNEL + +For example: a heterogeneous system with 1 AMD CPU is connected to +4 MI200 (Aldebaran) GPUs using xGMI. + +Some more heterogeneous hardware details: + +- The CPU UMC (Unified Memory Controller) is mostly the same as the GPU UMC. + They have chip selects (csrows) and channels. However, the layouts are different + for performance, physical layout, or other reasons. +- CPU UMCs use 1 channel, In this case UMC = EDAC channel. This follows the + marketing speak. CPU has X memory channels, etc. +- CPU UMCs use up to 4 chip selects, So UMC chip select = EDAC CSROW. +- GPU UMCs use 1 chip select, So UMC = EDAC CSROW. +- GPU UMCs use 8 channels, So UMC channel = EDAC channel. + +The EDAC subsystem provides a mechanism to handle AMD heterogeneous +systems by calling system specific ops for both CPUs and GPUs. + +AMD GPU nodes are enumerated in sequential order based on the PCI +hierarchy, and the first GPU node is assumed to have a Node ID value +following those of the CPU nodes after latter are fully populated:: + + $ ls /sys/devices/system/edac/mc/ + mc0 - CPU MC node 0 + mc1 | + mc2 |- GPU card[0] => node 0(mc1), node 1(mc2) + mc3 | + mc4 |- GPU card[1] => node 0(mc3), node 1(mc4) + mc5 | + mc6 |- GPU card[2] => node 0(mc5), node 1(mc6) + mc7 | + mc8 |- GPU card[3] => node 0(mc7), node 1(mc8) + +For example, a heterogeneous system with one AMD CPU is connected to +four MI200 (Aldebaran) GPUs using xGMI. This topology can be represented +via the following sysfs entries:: + + /sys/devices/system/edac/mc/.. + + CPU # CPU node + ├── mc 0 + + GPU Nodes are enumerated sequentially after CPU nodes have been populated + GPU card 1 # Each MI200 GPU has 2 nodes/mcs + ├── mc 1 # GPU node 0 == mc1, Each MC node has 4 UMCs/CSROWs + │   ├── csrow 0 # UMC 0 + │   │   ├── channel 0 # Each UMC has 8 channels + │   │   ├── channel 1 # size of each channel is 2 GB, so each UMC has 16 GB + │   │   ├── channel 2 + │   │   ├── channel 3 + │   │   ├── channel 4 + │   │   ├── channel 5 + │   │   ├── channel 6 + │   │   ├── channel 7 + │   ├── csrow 1 # UMC 1 + │   │   ├── channel 0 + │   │   ├── .. + │   │   ├── channel 7 + │   ├── .. .. + │   ├── csrow 3 # UMC 3 + │   │   ├── channel 0 + │   │   ├── .. + │   │   ├── channel 7 + │   ├── rank 0 + │   ├── .. .. + │   ├── rank 31 # total 32 ranks/dimms from 4 UMCs + ├ + ├── mc 2 # GPU node 1 == mc2 + │   ├── .. # each GPU has total 64 GB + + GPU card 2 + ├── mc 3 + │   ├── .. + ├── mc 4 + │   ├── .. + + GPU card 3 + ├── mc 5 + │   ├── .. + ├── mc 6 + │   ├── .. + + GPU card 4 + ├── mc 7 + │   ├── .. + ├── mc 8 + │   ├── ..