MLIR 24.0.0git
mlir::ROCDL::TargetInfo Class Reference

Describes the AMDGPU target a lowering is producing code for: the triple's subarch (which identifies the GPU) together with the resolved set of frontend-visible target features. More...

#include "mlir/Dialect/LLVMIR/ROCDLTargetInfo.h"

Public Types

using Feature = ::llvm::AMDGPU::AMDGPUFeature

Public Member Functions

 TargetInfo ()=default
 Constructs an unknown target: no subarch, and every feature query answers false.
bool has (Feature feature) const
 Returns whether the target has feature.
bool hasOcpFp8 () const
 Returns whether the target's fp8 conversions exist and use the OCP formats (E4M3FN/E5M2) rather than the FNUZ ones.
bool hasFnuzFp8 () const
 Returns whether the target has fp8 conversions that use the FNUZ formats (E4M3FNUZ/E5M2FNUZ).
bool isGeneration (unsigned major) const
 Returns whether the target belongs to gfx generation major (9 for any gfx9xx, 12 for any gfx12xx, ...).
std::optional< unsigned > getBufferResourceNumRecordsWidth () const
 Returns the width in bits of the num_records field of the buffer resource (V#), or nullopt for an unknown target.
std::optional< unsigned > getMaxAddressableLocalMemorySize () const
 Returns the maximum LDS in bytes a single workgroup can address, or nullopt for an unknown target.
std::optional< unsigned > getWavefrontSize () const
 Returns the wavefront size, or nullopt for an unknown target.
bool supportsBothWavefrontSizes () const
 Returns whether the GPU can be configured for 32-lane or 64-lane wavefronts.
std::optional< unsigned > getTotalNumSGPRs () const
 Returns the total number of SGPRs, or nullopt for an unknown target.
std::optional< unsigned > getAddressableNumSGPRs () const
 Returns the number of SGPRs addressable by a kernel, or nullopt for an unknown target.
std::optional< unsigned > getSGPRAllocGranule () const
 Returns the SGPR allocation granularity in registers, or nullopt for an unknown target.
std::optional< unsigned > getVGPRAllocGranule () const
 Returns the VGPR allocation granularity in registers, or nullopt for an unknown target.
std::optional< unsigned > getLDSBankCount () const
 Returns the number of LDS banks per compute unit, or nullopt for an unknown target.
std::optional< unsigned > getMaxWavesPerEU () const
 Returns the maximum number of waves per execution unit, ignoring any limits a particular kernel imposes, or nullopt for an unknown target.
::llvm::AMDGPU::TargetIDSetting getXnackSetting () const
 Returns whether xnack is on, off, either, or unsupported on this target.
::llvm::AMDGPU::TargetIDSetting getSramEccSetting () const
 Returns whether sramecc is on, off, either, or unsupported, as for getXnackSetting().
void migrateArchFeaturesToModuleFlags (Operation *op) const
 Records the xnack and sramecc settings this target's ID pinned onto the module op, as the rocdl.xnack and rocdl.sramecc attributes that translate to the amdgpu.xnack and amdgpu.sramecc module flags.
::llvm::AMDGPU::IsaVersion getIsaVersion () const
 Returns the ISA version.
::llvm::Triple::SubArchType getSubArch () const
::llvm::AMDGPU::GPUKind getGPUKind () const
StringRef getArchName () const
 Returns the canonical GPU name ("gfx942", "gfx9-4-generic"), or "" if the target is unknown.
bool isGeneric () const
 Returns whether this is a "gfxN-generic" target, which carries only the features common to every GPU it covers.
bool isUnknown () const
 Returns whether no GPU was identified, in which case every feature query answers false.
const ::llvm::AMDGPU::AMDGPUFeatureBitset & getFeatures () const

Static Public Member Functions

static FailureOr< TargetInfo > get (StringRef arch, unsigned waveSize=0, function_ref< InFlightDiagnostic()> emitError=nullptr)
 Resolves a target description.
static std::optional<::llvm::AMDGPU::TargetID > parseTargetID (StringRef arch)
 Parses arch into a target ID, accepting the spellings get() documents, or returns nullopt if it names no valid target.

Detailed Description

Describes the AMDGPU target a lowering is producing code for: the triple's subarch (which identifies the GPU) together with the resolved set of frontend-visible target features.

Lowerings should gate on features (has(FEAT_...)) rather than on ISA version arithmetic, and add features if necessary.

Definition at line 25 of file ROCDLTargetInfo.h.

Member Typedef Documentation

◆ Feature

using mlir::ROCDL::TargetInfo::Feature = ::llvm::AMDGPU::AMDGPUFeature

Definition at line 27 of file ROCDLTargetInfo.h.

Constructor & Destructor Documentation

◆ TargetInfo()

mlir::ROCDL::TargetInfo::TargetInfo ( )
default

Constructs an unknown target: no subarch, and every feature query answers false.

References mlir::emitError().

Referenced by get().

Member Function Documentation

◆ get()

FailureOr< TargetInfo > TargetInfo::get ( StringRef arch,
unsigned waveSize = 0,
function_ref< InFlightDiagnostic()> emitError = nullptr )
static

Resolves a target description.

arch names the architecture the way Clang does, and accepts any of:

  • a full target ID, "<triple>-<processor>[:<feature><+|->]*", such as "amdgcn-amd-amdhsa--gfx90a:sramecc+:xnack-" (what rocminfo prints for a device's ISA) or "amdgpu9.0a-amd-amdhsa--gfx90a";
  • a triple on its own, such as "amdgpu9.42-amd-amdhsa" or the legacy subarch-less "amdgcn-amd-amdhsa";
  • a processor on its own, with optional target-ID modifiers: "gfx942", "gfx942:xnack+", "gfx9-4-generic".

Only xnack and sramecc may be given as modifiers, and only on a processor that supports them; this is the same grammar clang::parseTargetID accepts, and it is validated by llvm::AMDGPU::TargetID.

waveSize pins the wavefront size for targets that run at either, and must be 0 (meaning the target's own default), 32, or 64.

Diagnostics are emitted via emitError.

Definition at line 81 of file ROCDLTargetInfo.cpp.

References mlir::emitError(), fail(), isUnknown(), parseTargetID(), resolveWavefrontSize(), and TargetInfo().

◆ getAddressableNumSGPRs()

std::optional< unsigned > TargetInfo::getAddressableNumSGPRs ( ) const

Returns the number of SGPRs addressable by a kernel, or nullopt for an unknown target.

This is below getTotalNumSGPRs() where some are reserved.

Definition at line 173 of file ROCDLTargetInfo.cpp.

References isUnknown().

◆ getArchName()

StringRef TargetInfo::getArchName ( ) const

Returns the canonical GPU name ("gfx942", "gfx9-4-generic"), or "" if the target is unknown.

Definition at line 236 of file ROCDLTargetInfo.cpp.

◆ getBufferResourceNumRecordsWidth()

std::optional< unsigned > TargetInfo::getBufferResourceNumRecordsWidth ( ) const

Returns the width in bits of the num_records field of the buffer resource (V#), or nullopt for an unknown target.

Definition at line 157 of file ROCDLTargetInfo.cpp.

◆ getFeatures()

const ::llvm::AMDGPU::AMDGPUFeatureBitset & mlir::ROCDL::TargetInfo::getFeatures ( ) const
inline

Definition at line 170 of file ROCDLTargetInfo.h.

◆ getGPUKind()

::llvm::AMDGPU::GPUKind mlir::ROCDL::TargetInfo::getGPUKind ( ) const
inline

Definition at line 156 of file ROCDLTargetInfo.h.

◆ getIsaVersion()

AMDGPU::IsaVersion TargetInfo::getIsaVersion ( ) const

Returns the ISA version.

For a generic target this is the floor of the family it covers (gfx9-4-generic reports 9.4.0), so it must not be used to decide whether an instruction is available.

Definition at line 232 of file ROCDLTargetInfo.cpp.

◆ getLDSBankCount()

std::optional< unsigned > TargetInfo::getLDSBankCount ( ) const

Returns the number of LDS banks per compute unit, or nullopt for an unknown target.

Definition at line 193 of file ROCDLTargetInfo.cpp.

References isUnknown().

◆ getMaxAddressableLocalMemorySize()

std::optional< unsigned > TargetInfo::getMaxAddressableLocalMemorySize ( ) const

Returns the maximum LDS in bytes a single workgroup can address, or nullopt for an unknown target.

Definition at line 161 of file ROCDLTargetInfo.cpp.

References isUnknown().

◆ getMaxWavesPerEU()

std::optional< unsigned > TargetInfo::getMaxWavesPerEU ( ) const

Returns the maximum number of waves per execution unit, ignoring any limits a particular kernel imposes, or nullopt for an unknown target.

Definition at line 199 of file ROCDLTargetInfo.cpp.

References isUnknown().

◆ getSGPRAllocGranule()

std::optional< unsigned > TargetInfo::getSGPRAllocGranule ( ) const

Returns the SGPR allocation granularity in registers, or nullopt for an unknown target.

Definition at line 179 of file ROCDLTargetInfo.cpp.

References isUnknown().

◆ getSramEccSetting()

::llvm::AMDGPU::TargetIDSetting mlir::ROCDL::TargetInfo::getSramEccSetting ( ) const
inline

Returns whether sramecc is on, off, either, or unsupported, as for getXnackSetting().

Definition at line 133 of file ROCDLTargetInfo.h.

◆ getSubArch()

::llvm::Triple::SubArchType mlir::ROCDL::TargetInfo::getSubArch ( ) const
inline

Definition at line 155 of file ROCDLTargetInfo.h.

◆ getTotalNumSGPRs()

std::optional< unsigned > TargetInfo::getTotalNumSGPRs ( ) const

Returns the total number of SGPRs, or nullopt for an unknown target.

Definition at line 167 of file ROCDLTargetInfo.cpp.

References isUnknown().

◆ getVGPRAllocGranule()

std::optional< unsigned > TargetInfo::getVGPRAllocGranule ( ) const

Returns the VGPR allocation granularity in registers, or nullopt for an unknown target.

This property is wavesize-dependent.

Definition at line 185 of file ROCDLTargetInfo.cpp.

References getWavefrontSize(), and isUnknown().

◆ getWavefrontSize()

std::optional< unsigned > TargetInfo::getWavefrontSize ( ) const

Returns the wavefront size, or nullopt for an unknown target.

Targets that support both sizes report 32 unless "+wavefrontsize64" was requested.

Definition at line 205 of file ROCDLTargetInfo.cpp.

References has().

Referenced by getVGPRAllocGranule().

◆ getXnackSetting()

::llvm::AMDGPU::TargetIDSetting mlir::ROCDL::TargetInfo::getXnackSetting ( ) const
inline

Returns whether xnack is on, off, either, or unsupported on this target.

"Any" means the target supports both and no :xnack+/- modifier was used.

Definition at line 127 of file ROCDLTargetInfo.h.

◆ has()

bool mlir::ROCDL::TargetInfo::has ( Feature feature) const
inline

Returns whether the target has feature.

Definition at line 64 of file ROCDLTargetInfo.h.

Referenced by getWavefrontSize(), hasFnuzFp8(), hasOcpFp8(), and isGeneration().

◆ hasFnuzFp8()

bool mlir::ROCDL::TargetInfo::hasFnuzFp8 ( ) const
inline

Returns whether the target has fp8 conversions that use the FNUZ formats (E4M3FNUZ/E5M2FNUZ).

Definition at line 74 of file ROCDLTargetInfo.h.

References has(), and hasOcpFp8().

◆ hasOcpFp8()

bool mlir::ROCDL::TargetInfo::hasOcpFp8 ( ) const
inline

Returns whether the target's fp8 conversions exist and use the OCP formats (E4M3FN/E5M2) rather than the FNUZ ones.

Definition at line 68 of file ROCDLTargetInfo.h.

References has().

Referenced by hasFnuzFp8().

◆ isGeneration()

bool TargetInfo::isGeneration ( unsigned major) const

Returns whether the target belongs to gfx generation major (9 for any gfx9xx, 12 for any gfx12xx, ...).

Prefer has() where a feature expresses the condition; this is used when no feature exists and the property being checked is a function of the major ISA generation (such as the details of buffer encoding).

Definition at line 123 of file ROCDLTargetInfo.cpp.

References has(), and isUnknown().

◆ isGeneric()

bool TargetInfo::isGeneric ( ) const

Returns whether this is a "gfxN-generic" target, which carries only the features common to every GPU it covers.

Definition at line 240 of file ROCDLTargetInfo.cpp.

References isUnknown().

◆ isUnknown()

bool mlir::ROCDL::TargetInfo::isUnknown ( ) const
inline

Returns whether no GPU was identified, in which case every feature query answers false.

Definition at line 168 of file ROCDLTargetInfo.h.

Referenced by get(), getAddressableNumSGPRs(), getLDSBankCount(), getMaxAddressableLocalMemorySize(), getMaxWavesPerEU(), getSGPRAllocGranule(), getTotalNumSGPRs(), getVGPRAllocGranule(), isGeneration(), and isGeneric().

◆ migrateArchFeaturesToModuleFlags()

void TargetInfo::migrateArchFeaturesToModuleFlags ( Operation * op) const

Records the xnack and sramecc settings this target's ID pinned onto the module op, as the rocdl.xnack and rocdl.sramecc attributes that translate to the amdgpu.xnack and amdgpu.sramecc module flags.

These flags are given as :{xnack,sramecc} target-ID "modifiers", since they used to be subtarget features, but now frontends (like us and Clang) need to migrate them into module flags. This representation keeps us compatible with Clang and the output of tools like rocminfo.

If a particular modifier is not given, no attribute is set for it, putting that value into its "any" state if it is controllable.

Definition at line 213 of file ROCDLTargetInfo.cpp.

References mlir::Operation::getContext(), mlir::MLIRContext::getOrLoadDialect(), and mlir::LLVM::satisfiesLLVMModule().

◆ parseTargetID()

std::optional< AMDGPU::TargetID > TargetInfo::parseTargetID ( StringRef arch)
static

Parses arch into a target ID, accepting the spellings get() documents, or returns nullopt if it names no valid target.

Use only if you need to get the individual components of the target ID.

Definition at line 32 of file ROCDLTargetInfo.cpp.

Referenced by get().

◆ supportsBothWavefrontSizes()

bool mlir::ROCDL::TargetInfo::supportsBothWavefrontSizes ( ) const
inline

Returns whether the GPU can be configured for 32-lane or 64-lane wavefronts.

Definition at line 100 of file ROCDLTargetInfo.h.


The documentation for this class was generated from the following files: