Skip to content

Latest commit

 

History

59 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

NVMeNext

Introduction

NVMeNext is a scalable software-defined virtual NVMe device for GPU-centric storage systems. Built on NVMeVirt, it is implemented as a Linux kernel module that emulates an NVMe SSD at the PCI layer, so the device is visible to GPUs and supports kernel bypassing and PCI peer-to-peer DMA out of the box. NVMeNext rearchitects the I/O path with a multi-manager design, a NAND performance model with per-plane fine-grained locking, and DSA-accelerated batched request processing, enabling it to sustain MIOPS-scale fine-grained random reads issued directly by GPU threads. Unmodified GPU-initiated I/O frameworks such as BaM run on NVMeNext without any changes.

For our existing emulator NVMeVirt, which supports conventional SSDs, NVM SSDs, ZNS SSDs, and more, please refer to this link.

Installation

System requirements

NVMeNext requires a platform with Intel Data Streaming Accelerator (DSA). Make sure that your processor supports DSA, and that both DSA and VT-d (Intel IOMMU) are enabled in the BIOS.

The recommended Linux kernel version is v6.8 and higher (tested on Linux vanilla kernel v6.8.0). The kernel must be built with the Intel Data Accelerator Driver (CONFIG_INTEL_IDXD) and its shared virtual memory support (CONFIG_INTEL_IDXD_SVM=y) enabled.

Setting kernel parameters

A part of the main memory should be reserved for the storage of the emulated NVMe device. To reserve a chunk of physical memory, add the following option to GRUB_CMDLINE_LINUX in /etc/default/grub as follows:

GRUB_CMDLINE_LINUX="memmap=64G\\\$128G"

This example will reserve 64GiB of physical memory chunk (out of the total 192GiB physical memory) starting from the 128GiB memory offset. You may need to adjust those values depending on the available physical memory size and the desired storage capacity.

In addition, NVMeNext uses the shared work queues (WQs) of the Intel DSA engine for data movement. Shared WQs require the IOMMU to be enabled in scalable mode, so the following options must also be added to the kernel command line:

GRUB_CMDLINE_LINUX="memmap=64G\\\$128G iommu=on intel_iommu=on,sm_on"

After changing the /etc/default/grub file, you are required to run the following commands to update grub and reboot your system.

$ sudo update-grub
$ sudo reboot

Setting up DSA

NVMeNext uses accel-config to configure DSA devices and work queues. Please install it by following the instructions in the idxd-config repository.

Once accel-config is installed, run the provided script to enable the DSA devices, engines, and work queues (WQs):

$ ./setup_dsa.sh init

The script scans all DSA devices in the system and, for each device, configures its engines and WQs as a single group, with all WQs in shared mode.

Compiling nvmenext

Please download the latest version of nvmenext from Github:

$ git clone https://github.com/snu-csl/nvmenext

nvmenext is implemented as a Linux kernel module. Thus, the kernel headers should be installed in the /lib/modules/$(shell uname -r) directory to compile nvmenext.

You may find the detailed SSD configuration parameters from ssd_config.h.

Build the kernel module by running the make command in the nvmenext source directory.

$ make
make -C /lib/modules/6.8.0/build M=/path/to/nvmenext modules
make[1]: Entering directory '/path/to/linux-headers-6.8.0'
  CC [M]  /path/to/nvmenext/main.o
  CC [M]  /path/to/nvmenext/pci.o
  CC [M]  /path/to/nvmenext/admin.o
  CC [M]  /path/to/nvmenext/io.o
  CC [M]  /path/to/nvmenext/dma.o
  CC [M]  /path/to/nvmenext/dsa.o
  CC [M]  /path/to/nvmenext/ssd.o
  CC [M]  /path/to/nvmenext/conv_ftl.o
  CC [M]  /path/to/nvmenext/pqueue/pqueue.o
  CC [M]  /path/to/nvmenext/channel_model.o
  LD [M]  /path/to/nvmenext/nvmenext.o
  MODPOST /path/to/nvmenext/Module.symvers
  CC [M]  /path/to/nvmenext/nvmenext.mod.o
  LD [M]  /path/to/nvmenext/nvmenext.ko
  BTF [M] /path/to/nvmenext/nvmenext.ko
make[1]: Leaving directory '/path/to/linux-headers-6.8.0'
$

Using nvmenext

nvmenext is configured to emulate the Conventional SSD by default. You can attach an emulated Conventional SSD in your system by loading the nvmenext kernel module as follows:

$ sudo insmod ./nvmenext.ko \
  memmap_start=128G         \
  memmap_size=64G           \
  gpu_start=0x203000000000  \
  gpu_size=4G               \
  dispatcher_cpus=0,1,2,...,15,32,33,34,...,47 \
  worker_cpus=16

In the above example, memmap_start and memmap_size indicate the relative offset and the size of the reserved memory, respectively. Those values should match the configurations specified in the /etc/default/grub file shown earlier. gpu_start and gpu_size indicate the physical address range assigned to the GPU memory (i.e., BAR1 of the GPU). You can find these values by reading the resource file of the GPU device in sysfs, where the second line corresponds to BAR1:

$ cat /sys/bus/pci/devices/<BDF>/resource

In addition, dispatcher_cpus specifies the ids of cores on which the request manager threads run. Each manager handles the entire life cycle of an NVMe command, i.e., fetching commands from the SQ, computing the target completion time, and posting completions to the CQ. Assigning multiple cores allows NVMeNext to process requests in parallel across managers. worker_cpus specifies the core for the worker thread. Although the worker thread does not play a major role in NVMeNext, one worker thread is required for the emulator to operate correctly.

It is highly recommended to use the isolcpus Linux command-line configuration to avoid schedulers putting tasks on the CPUs that NVMeNext uses:

GRUB_CMDLINE_LINUX="memmap=64G\\\$128G iommu=on intel_iommu=on,sm_on isolcpus=0,1,2,...,15,16,32,33,34,...,47"

Currently, NVMeNext assumes the following configuration: 32 manager threads and one worker thread, a single DSA device with 8 shared WQs, and a static mapping in which every four manager threads share one WQ. If you want to change the DSA configuration, such as using more DSA devices or changing the manager-to-WQ mapping, you need to modify the device initialization and deinitialization routines in dsa.c.

When you are successfully load the nvmenext module, you can see something like these from the system message.

$ sudo dmesg
[  867.928223] NVMeNext: Storage: 0x1000100000-0x3000000000 (131071 MiB)
[  868.083446] PCI host bridge to bus 0001:10
[  868.083450] pci_bus 0001:10: root bus resource [io  0x0000-0xffff]
...
(log omitted)
...
[  870.790429] nvme nvme0: pci function 0001:10:00.0
[  870.792752] nvme nvme0: 1/0/0 default/read/poll queues
[  870.795793] NVMeNext: Virtual NVMe device created

Now the emulated nvmenext device is ready to be used as shown below. The actual device number (/dev/nvme0) can vary depending on the number of real NVMe devices in your system.

$ ls -l /dev/nvme*
crw------- 1 root root 242, 0 Feb 22 14:13 /dev/nvme0
brw-rw---- 1 root disk 259, 5 Feb 22 14:13 /dev/nvme0n1

Profiling

To use nvmenext as an I/O trace profiler, see io-profiler for detailed instructions.

License

NVMeNext is offered under the terms of the GNU General Public License version 2 as published by the Free Software Foundation. More information about this license can be found here.

Priority queue implementation pqueue/ is offered under the terms of the BSD 2-clause license (GPL-compatible). (Copyright (c) 2014, Volkan Yazıcı volkan.yazici@gmail.com. All rights reserved.)

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages