Intel Alder Lake-S Architecture




 12th Gen Intel® Core™ CPUs adapt to the ways you work and play. When gaming, the processor prevents background tasks from interrupting or using your high-performance cores. When working, it provides a smoother system-level experience while using demanding applications.


12Th Generation Processors Line-up's [ Intel Alder Lake-S Architecture ]



Highlighted :


12th Gen CPUs integrate two types of cores into a single die: performance-cores (P-cores) and efficient-cores (E-cores).

Performance-cores :


  • Physically larger high-performance cores designed for raw speed while maintaining efficiency.
  • Optimized for low-latency single-threaded performance and AI workloads.
  • Capable of hyper-threading, or running two software threads at once.
  • Measured at 19% better performance, on average, than 11th Gen Intel® Core™ CPUs across a wide range of workloads at ISO frequency3.


Efficient-cores :


  • Physically smaller, with multiple E-cores fitting into the physical space occupied by one P-core.
  • Optimized for multi-core performance-per-watt—delivering scalable multithread performance and efficient offload of background tasks.
  • Capable of running a single software thread.
  • Capable of 40% more performance when running at the same power as a single Skylake core4.




DDR5 Memory Details :


DDR5 is the next-generation specification for RAM and it comes with a host of improvements in speed and efficiency when compared to DDR4, the current standard.

  • Higher-bandwidth kits thanks to doubled burst length—the number of bits that can be read per cycle.
  • 12th Gen supports speeds up to 4,800MHz for DDR5 and 3,200MHz for DDR4.
  • DDR5 allows capacities of up to 128GB of RAM per module, whereas DDR4 allows only 32GB.
  • DDR5 doubles the number of memory bank groups and improves the speed at which groups can be refreshed.



PCIe 5.0 :


12th Gen Intel® Core™ CPUs are at the forefront of the industry transition to PCIe 5.0. PCIe 5.0 doubles the bandwidth of 4.0, which means your system will be ready for the next generation of SSDs and discrete GPUs.

PCIe is the high-bandwidth expansion bus used to connect graphics cards, SSDs, and other peripherals to your motherboard. Each generation of PCIe doubles in throughput, with PCIe 5.0 providing theoretical maximum data transfer speeds of 32 GT/s.

  • Full backwards compatibility with PCIe 4.0 and 3.0 devices.
  • Double the bandwidth of 4.0 and four times the bandwidth of 3.0.
  • Up to 16 CPU PCIe 5.0 lanes and up to 4 CPU PCIe 4.0 lanes.



Intel UHD Graphics 64EU :


The UHD Graphics 64EU is a mobile integrated graphics solution by Intel, launched on January 4th, 2022. Built on the 10 nm process, and based on the Alder Lake GT1 graphics processor, the device supports DirectX 12. This ensures that all modern games will run on UHD Graphics 64EU. It features 512 shading units, 32 texture mapping units, and 16 ROPs. The GPU is operating at a frequency of 300 MHz, which can be boosted up to 1400 MHz.
Its power draw is rated at 45 W maximum.


General info


Of UHD Graphics 64EU's architecture, market segment and release date.

  • Place in performance rating not rated
  • Architecture Generation 12.2 (2021−2022)
  • GPU code name Alder Lake GT1
  • Market segment Desktop
  • Release date 4 January 2022 (less than a year ago)

Technical specs


  • Pipelines / CUDA cores 512 of 18432 (AD102)
  • Boost clock speed 1400 MHz of 2903 (Radeon Pro W6600)
  • Manufacturing process technology 10 nm of 4 (H100 PCIe)
  • Thermal design power (TDP) 45 Watt of 900 (Tesla S2050)
  • Texture fill rate 44.80 of 939.8 (H100 SXM5)

Memory :


  • Memory type System Shared
  • Maximum RAM amount System Shared of 128 (Radeon Instinct MI250X)
  • Memory bus width System Shared of 8192 (Radeon Instinct MI250X)
  • Memory clock speed System Shared of 21000 (GeForce RTX 3090 Ti)

API support :


DirectX 12 (12_1)
Shader Model 6.4
OpenGL 4.6
OpenCL 3.0
Vulkan 1.3





Nvidia Ampere Architecture

 



Ampere is the codename for a graphics processing unit (GPU) microarchitecture developed by Nvidia as the successor to both the Volta and Turing architectures, officially announced on May 14, 2020. It is named after French mathematician and physicist André-Marie Ampère. Nvidia announced the next-generation GeForce 30 series consumer GPUs at a GeForce Special Event on September 1, 2020. Nvidia announced 80GB GPU at SC20 on November 16, 2020. Mobile RTX graphics cards and the RTX 3060 were revealed on January 12, 2021. Nvidia also announced Ampere's successor, Hopper, at GTC 2022, and "Ampere Next Next" for a 2024 release at GPU Technology Conference 2021.


Ampere Graphics Processors Line-up's 

Highlighted :


Third-Generation Tensor Cores


First introduced in the NVIDIA Volta™ architecture, NVIDIA Tensor Core technology has brought dramatic speedups to AI, bringing down training times from weeks to hours and providing massive acceleration to inference. The NVIDIA Ampere architecture builds upon these innovations by bringing new precisions—Tensor Float 32 (TF32) and floating point 64 (FP64)—to accelerate and simplify AI adoption and extend the power of Tensor Cores to HPC.

TF32 works just like FP32 while delivering speedups of up to 20X for AI without requiring any code change. Using NVIDIA Automatic Mixed Precision, researchers can gain an additional 2X performance with automatic mixed precision and FP16 by adding just a couple of lines of code. And with support for bfloat16, INT8, and INT4, Tensor Cores in NVIDIA Ampere architecture Tensor Core GPUs create an incredibly versatile accelerator for both AI training and inference. Bringing the power of Tensor Cores to HPC, A100 and A30 GPUs also enable matrix operations in full, IEEE-certified, FP64 precision.


Third-Generation NVLink


Scaling applications across multiple GPUs requires extremely fast movement of data. The third generation of NVIDIA® NVLink® in the NVIDIA Ampere architecture doubles the GPU-to-GPU direct bandwidth to 600 gigabytes per second (GB/s), almost 10X higher than PCIe Gen4. When paired with the latest generation of NVIDIA NVSwitch™, all GPUs in the server can talk to each other at full NVLink speed for incredibly fast data transfers. 

NVIDIA DGX™A100 and servers from other leading computer makers take advantage of NVLink and NVSwitch technology via NVIDIA HGX™ A100 baseboards to deliver greater scalability for HPC and AI workloads.



Second-Generation RT Cores


The NVIDIA Ampere architecture’s second-generation RT Cores in the NVIDIA A40 deliver massive speedups for workloads like photorealistic rendering of movie content, architectural design evaluations, and virtual prototyping of product designs. RT Cores also speed up the rendering of ray-traced motion blur for faster results with greater visual accuracy and can simultaneously run ray tracing with either shading or denoising capabilities.




Architectural improvements of the Ampere architecture include the following:


  • CUDA Compute Capability 8.0 for A100 and 8.6 for the GeForce 30 series
  • TSMC's 7 nm FinFET process for A100
  • Custom version of Samsung's 8 nm process (8N) for the GeForce 30 series
  • Third-generation Tensor Cores with FP16, bfloat16, TensorFloat-32 (TF32) and FP64 support and sparsity acceleration. The individual Tensor cores have with 256 FP16 FMA operations per second 4x processing power (GA100 only, 2x on GA10x) compared to previous Tensor Core generations; the Tensor Core Count is reduced to one per SM.
  • Second-generation ray tracing cores; concurrent ray tracing, shading, and compute for the GeForce 30 series
  • High Bandwidth Memory 2 (HBM2) on A100 40GB & A100 80GB
  • GDDR6X memory for GeForce RTX 3090, RTX 3080 Ti, RTX 3080, RTX 3070 Ti
  • Double FP32 cores per SM on GA10x GPUs
  • NVLink 3.0 with a 50Gbit/s per pair throughput
  • PCI Express 4.0 with SR-IOV support (SR-IOV is reserved only for A100)
  • Multi-instance GPU (MIG) virtualization and GPU partitioning feature in A100 supporting up to seven instances
  • PureVideo feature set K hardware video decoding with AV1 hardware decoding for the GeForce 30 series and feature set J for A100
  • 5 NVDEC for A100
  • Adds new hardware-based 5-core JPEG decode (NVJPG) with YUV420, YUV422, YUV444, YUV400, RGBA. Should not be confused with Nvidia NVJPEG (GPU-accelerated library for JPEG encoding/decoding)




Ampere PowerFull GPU :




Nvidia GeForce RTX 3090

The GeForce RTX 3090 Ti is anenthusiast-class graphics card by NVIDIA, launched on January 27th, 2022. Built on the 8 nm process, and based on the GA102 graphics processor, in its GA102-350-A1 variant, the card supports DirectX 12 Ultimate. This ensures that all modern games will run on GeForce RTX 3090 Ti. Additionally, the DirectX 12 Ultimate capability guarantees support for hardware-raytracing, variable-rate shading and more, in upcoming video games. The GA102 graphics processor is a large chip with a die area of 628 mm² and 28,300 million transistors. It features 10752 shading units, 336 texture mapping units, and 112 ROPs. Also included are 336 tensor cores which help improve the speed of machine learning applications. The card also has 84 raytracing acceleration cores. NVIDIA has paired 24 GB GDDR6X memory with the GeForce RTX 3090 Ti, which are connected using a 384-bit memory interface. The GPU is operating at a frequency of 1560 MHz, which can be boosted up to 1860 MHz, memory is running at 1313 MHz (21 Gbps effective).

Being a triple-slot card, the NVIDIA GeForce RTX 3090 Ti draws power from 1x 16-pin power connector, with power draw rated at 450 W maximum. Display outputs include: 1x HDMI 2.1, 3x DisplayPort 1.4a. GeForce RTX 3090 Ti is connected to the rest of the system using a PCI-Express 4.0 x16 interface. The card's dimensions are 336 mm x 140 mm x 61 mm, and it features a triple-slot cooling solution. Its price at launch was 1999 US Dollars.



Ampere Vs Turing Architecture


The fastest RTX graphics cards are now alive, from Nvidia’s factory. New Nvidia Ampere GPUs, the successor of Turing are most powerful, that’s what we expect from the new-gen. Specifically, ray tracing performance has improved so much.

The Turing architecture also introduced Ray Tracing cores used to accelerate photo realistic rendering. With Ampere NVIDIA has continued to make significant improvements










Nvidia Introduction

Introduction Of Nvidia :


Nvidia Corporation commonly known as Nvidia, is an American multinational technology company incorporated in Delaware and based in Santa Clara, California. It is a software and fabless company which designs graphics processing units (GPUs), application programming interface (APIs) for data science and high-performance computing as well as system on a chip units (SoCs) for the mobile computing and automotive market. Nvidia is a global leader in artificial intelligence hardware and software. Its professional line of GPUs are used in workstations for applications in such fields as architecture, engineering and construction, media and entertainment, automotive, scientific research, and manufacturing design.

In addition to GPU manufacturing, Nvidia provides an API called CUDA that allows the creation of massively parallel programs which utilize GPUs. They are deployed in supercomputing sites around the world. More recently, it has moved into the mobile computing market, where it produces Tegra mobile processors for smartphones and tablets as well as vehicle navigation and entertainment systems.In addition to AMD, its competitors include Intel, Qualcomm and AI-accelerator companies such as Graphcore.



Architectures :





Features 


Geforce Model's :



Nvidia RTX :


NVIDIA RTX technology empowers developers to redefine what's possible in computer graphics, video, and imaging. Accelerate application development by leveraging the powerful new ray tracing, deep learning, and rasterization capabilities through industry-leading software Platforms, SDKs and APIs.



Nvidia GTX :


GTX stands for Giga Texel Shader eXtreme and is a variant under the brand GeForce owned by Nvidia. They were first introduced in 2008 with series 200, codenamed Tesla. The first product in this series was GTX 260 and more expensive GTX 280. The introduction of these cards also affected the naming scheme and from the release of these cards onwards, Nvidia GPUs used a naming scheme that has GTX/GT as a prefix followed by their model number. With every other major release in the series, Nvidia changed its microarchitecture on which its cards are based on i.e. series 200 & 300 were based on Tesla architecture, series 400 & 500 were based on Fermi architecture and so on.
The latest GTX series 16, consist of GTX 1650, GTX 1660, GTX 1660Ti, and its Super counterparts. These are based on Turing architecture and were introduced in 2019.


Nvidia GTS :


Built from the ground up for next generation DX11 gaming, the GeForce GTS 450 delivers revolutionary tessellation performance for the ultimate gaming experience. With full support for NVIDIA 3D Vision the GeForce GTS 450 provides the graphics horsepower and video bandwidth needed to experience games and high definition Blu-ray movies in eye-popping stereoscopic 3D.



Nvidia GT :


The Gigabyte GeForce GT 1030 is one of the best entry-level GPUs. With its ultra-durable components, this GPU offers outstanding performance without compromising the system's lifespan. If you are a gaming enthusiast, you will love this GPU.






The First Graphics Processor Of Nvidia :


  • GeForce 256


The term GPU has been in use since at least the 1980s. Nvidia popularized it in 1999 by marketing the GeForce 256 add-in board (AIB) as the world’s first GPU. It offered integrated transform, lighting, triangle setup/clipping, and rendering engines as a single-chip processor.

Very-large-scale integrated circuitry—VLSI, started taking hold in the early 1990s. As the number of transistors engineers could incorporate on a single chip increased almost exponentially, the number of functions in the CPU and the graphics processor increased. One of the biggest consumers of the CPU was graphics transformation compute elements into graphics processors. Architects from various graphics chip companies decided transform and lighting (T&L) was a function that should be in the graphics processor. The operation was known at the time as transform and lighting (T&L). A T&L engine is a vertex shader and a geometry translator—many names for the little FFP.





Geforce 256 Specifications :


The GeForce 256 SDR was a graphics card by NVIDIA, launched on October 11th, 1999. Built on the 220 nm process, and based on the NV10 graphics processor, the card supports DirectX 7.0. Since GeForce 256 SDR does not support DirectX 11 or DirectX 12, it might not be able to run all the latest games. The NV10 graphics processor is an average sized chip with a die area of 139 mm² and 17 million transistors. It features 4 pixel shaders and 0 vertex shaders, 4 texture mapping units, and 4 ROPs. Due to the lack of unified shaders you will not be able to run recent games at all (which require unified shader/DX10+ support). NVIDIA has paired 32 MB SDR memory with the GeForce 256 SDR, which are connected using a 64-bit memory interface. The GPU is operating at a frequency of 120 MHz, memory is running at 143 MHz.
Being a single-slot card, the NVIDIA GeForce 256 SDR does not require any additional power connector, its power draw is not exactly known. Display outputs include: 1x VGA. GeForce 256 SDR is connected to the rest of the system using an AGP 4x interface.





Nvidia Success Story :


Nvidia was founded in 1993 by Jensen Huang, Chris Malachowsky, and Curtis Priem, the same year the term “millennial” was coined. Is this a “millennial” company? All signs point to yes as Nvidia was started with a belief that a PC would become a commercial device for enjoying video games and multimedia. What you have right now is the more advanced version of the chunky display device, a noisy CPU, clunky keyboard, and a ball mouse – all once called a PC, personal computer. At the time when the company started, there were several graphics chips companies, a number that soon multiplied manifold three years later.


With grit and determination, three young electrical engineers started Nvidia to make advanced specialized chips that would create faster and realistic graphics for video games. “There was no market in 1993, but we saw a wave coming,” said Malachowsky to Forbes. “There’s a California surfing competition that happens in a five-month window every year. When they see some type of wave phenomenon or storm in Japan, they tell all the surfers to show up in California, because there’s going to be a wave in two days. That’s what it was. We were at the beginning.”




Nvidia Fermi Architecture

 



Fermi is the codename for a graphics processing unit (GPU) microarchitecture developed by Nvidia, first released to retail in April 2010, as the successor to the Tesla microarchitecture. It was the primary microarchitecture used in the GeForce 400 series and GeForce 500 series. It was followed by Kepler, and used alongside Kepler in the GeForce 600 series, GeForce 700 series, and GeForce 800 series, in the latter two only in mobile GPUs. In the workstation market, Fermi found use in the Quadro x000 series, Quadro NVS models, as well as in Nvidia Tesla computing modules. All desktop Fermi GPUs were manufactured in 40nm, mobile Fermi GPUs in 40nm and 28nm. Fermi is the oldest microarchitecture from NVIDIA that received support for the Microsoft's rendering API Direct3D 12 feature_level 11.



Fermi Graphics Processors Line-up's 



Highlighted :


Fermi Graphic Processing Units (GPUs) feature 3.0 billion transistors and a schematic is sketched in Fig. 1

.

  • Streaming Multiprocessor (SM): composed of 32 CUDA cores (see Streaming Multiprocessor and CUDA core sections).
  • GigaThread global scheduler: distributes thread blocks to SM thread schedulers and manages the context switches between threads during execution (see Warp Scheduling section).
  • Host interface: connects the GPU to the CPU via a PCI-Express v2 bus (peak transfer rate of 8GB/s).
  • DRAM: supported up to 6GB of GDDR5 DRAM memory thanks to the 64-bit addressing capability (see Memory Architecture section).
  • Clock frequency: 1.5 GHz (not released by NVIDIA, but estimated by Insight 64).
  • Peak performance: 1.5 TFlops.
  • Global memory clock: 2 GHz.
  • DRAM bandwidth: 192GB/s.



Fermi Chips :


  • GF 100
  • GF 104
  • GF 106
  • GF 108
  • GF 110
  • GF 114
  • GF 116
  • GF 118
  • GF 119
  • GF 117


Architecture :



                                          




With these requests in mind, the Fermi team designed a processor that greatly increases raw compute horsepower, and through architectural innovations, also offers dramatically increased programmability and compute efficiency. The key architectural highlights of Fermi are:

• Third Generation Streaming Multiprocessor (SM)   

o 32 CUDA cores per SM, 4x over GT200 

o 8x the peak double precision floating point performance over GT200 

o Dual Warp Scheduler simultaneously schedules and dispatches instructions from two independent warps 

o 64 KB of RAM with a configurable partitioning of shared memory and L1 cache 


• Second Generation Parallel Thread Execution ISA 


o Unified Address Space with Full C++ Support 

o Optimized for OpenCL and DirectCompute
 
o Full IEEE 754-2008 32-bit and 64-bit precision 

o Full 32-bit integer path with 64-bit extensions 

o Memory access instructions to support transition to 64-bit addressing 

o Improved Performance through Predication 


• Improved Memory Subsystem 


o NVIDIA Parallel DataCacheTM hierarchy with Configurable L1 and Unified L2 Caches 

o First GPU with ECC memory support 

o Greatly improved atomic memory operation performance


• NVIDIA GigaThreadTM Engine


o 10x faster application context switching 

o Concurrent kernel execution o Out of Order thread block execution 

o Dual overlapped memory transfer engines 



More Details :

Optimized for OpenCL and DirectCompute 


OpenCL and DirectCompute are closely related to the CUDA programming model, sharing the key abstractions of threads, thread blocks, grids of thread blocks, barrier synchronization, perblock shared memory, global memory, and atomic operations. Fermi, a third-generation CUDA architecture, is by nature well-optimized for these APIs. In addition, Fermi offers hardware support for OpenCL and DirectCompute surface instructions with format conversion, allowing graphics and compute programs to easily operate on the same data. The PTX 2.0 ISA also adds support for the DirectCompute instructions population count, append, and bit-reverse. 







Fermi's PowerFull GPU :


  • NVIDIA GeForce GTX 590


The GeForce GTX 590 was an enthusiast-class graphics card by NVIDIA, launched on March 24th, 2011. Built on the 40 nm process, and based on the GF110 graphics processor, in its GF110-351-A1 variant, the card supports DirectX 12. Even though it supports DirectX 12, the feature level is only 11_0, which can be problematic with newer DirectX 12 titles. The GF110 graphics processor is a large chip with a die area of 520 mm² and 3,000 million transistors. GeForce GTX 590 combines two graphics processors to increase performance. It features 512 shading units, 64 texture mapping units, and 48 ROPs, per GPU. NVIDIA has paired 3,072 MB GDDR5 memory with the GeForce GTX 590, which are connected using a 384-bit memory interface per GPU (each GPU manages 1,536 MB). The GPU is operating at a frequency of 608 MHz, memory is running at 854 MHz (3.4 Gbps effective).






Intel Clarkdale Architecture

 Intel Clarkdale Architecture :




Clarkdale is theClarkdale is the codename for Intel's first-generation Core i5, i3 and Pentium dual-core desktop processors. It is closely related to the mobile Arrandale processor; both use dual-core dies based on the 32 nm Westmere microarchitecture and have integrated Graphics, PCI Express and DMI links built-in. codename for Intel's first-generation Core i5, i3 and Pentium dual-core desktop processors. It is closely related to the mobile Arrandale processor; both use dual-core dies based on the 32 nm Westmere microarchitecture and have integrated Graphics, PCI Express and DMI links built-in.



Clarkdale Processors Line-up's 




More Details


Four chipsets (all using the LGA 1156 socket) are compatible with the Clarkdale platform: the H55, H57, Q57, and standard Lynnfield P55-based motherboards. Here’s where it gets interesting. H55, H57, and Q57-based boards are identical in their overall construction, with each offering a new subset of Intel features as you go up the price range. H57-based motherboards can support two additional USB ports, two extra PCI Express x1 lanes, and support for Intel’s RAID-based Rapid Storage Technology. Q57 boards, more for business use, include Intel’s Active Management Technology—remote technical support. You can stick a Clarkdale processor in a P55 motherboard or, vice versa, a Lynnfield processor in an H55, H57, or Q57 motherboard. Either situation forces you to use a discrete graphics card, however.

Graphics Details


As mentioned, integrated gaming performance isn’t for tough titles. While Clarkdale systems might thrive on less demanding titles, the CPU’s integrated graphics weren’t enough to deliver playable frame-rates on PC World’s Unreal Tournament 3 benchmark at anything but a 1024-by-768 resolution screen at medium quality settings or less. And a forewarning: the sixteen PCI Express x16 lanes supported by Clarkdale chips cannot be split into dual x8 lanes for CrossFire or SLI should you aspire to transform your Clarkdale rig into a souped-up gaming machine. Clarkdale intends to make its mark on more common computers… including those in your living room.