Bajo nivel · C++Low-level · C++

Anatomía de un driver de Windows: del DriverEntry al IRPAnatomy of a Windows driver: from DriverEntry to the IRP

En esta páginaOn this page
  1. El pantallazo más común está mal entendido
  2. Un driver no es un programa
  3. IRQL, la ley física del kernel
  4. El viaje de un IRP
  5. Compruébalo desde modo usuario
  6. La tabla de dispatch
  7. IOCTL propios
  8. Sincronización: ¿encuentras el bug?
  9. Ciclo de vida: crear y descargar
  10. Cómo se trabaja de verdad
  11. En el kernel no hay red
  1. The most common crash is misunderstood
  2. A driver isn't a program
  3. IRQL, the kernel's law of physics
  4. An IRP's journey
  5. Check it from user mode
  6. The dispatch table
  7. Custom IOCTLs
  8. Synchronization: can you spot the bug?
  9. Lifecycle: create and unload
  10. How it's really done
  11. There's no safety net in the kernel

El pantallazo más común está mal entendido#

Si has visto alguna pantalla azul, es probable que dijera IRQL_NOT_LESS_OR_EQUAL. Es el bugcheck 0xA, y su nombre invita a pensar que alguien rompió una regla abstracta sobre niveles de interrupción. Casi nunca es eso.

En la gran mayoría de casos reales, detrás hay algo mucho más prosaico: un puntero malo, NULL o que apunta a memoria ya liberada, o memoria paginable tocada a un IRQL demasiado alto para que el fallo de página se pueda resolver. A DISPATCH_LEVEL o por encima, el sistema no puede traer páginas del disco. Lo que en un programa normal sería un fallo de página rutinario, que el sistema resuelve sin que te enteres, en el kernel se convierte en un bugcheck. Su pariente, el 0xD1 (DRIVER_IRQL_NOT_LESS_OR_EQUAL), va un paso más allá y señala directamente a un driver.

Por eso, cuando depuras uno de estos, la pregunta útil no es «¿qué regla de IRQL he roto?», sino esta:

La pregunta que hay que hacerse

¿Qué puntero he desreferenciado? ¿Era válido? ¿Y estaba garantizado que su memoria estuviera presente, dado el IRQL al que corría mi código?

En De Ring 3 a Ring 0 contamos la frontera entre tus programas y el kernel desde fuera: cómo se cruza y por qué un fallo al otro lado tumba el sistema entero. Hoy cruzamos esa frontera y miramos cómo está organizado el código que vive dentro: un driver de Windows, desde que se carga hasta que atiende una petición. Para entender esa pregunta, primero hay que entender qué es un driver y qué es el IRQL.

Un driver no es un programa#

Un ejecutable es pasivo: el sistema operativo lo arranca, le da su propio espacio de memoria y, si falla, cierra el proceso y limpia lo que dejó. Un driver no tiene nada de eso. Corre como parte del sistema operativo, en Ring 0, y no hay ninguna frontera de proceso que proteja al resto del sistema de sus errores.

Tampoco tiene main. Su punto de entrada es DriverEntry, y lo pone en marcha el I/O Manager de Windows al cargar el driver, sin la cadena de arranque de la biblioteca de C que en un programa normal acaba llamando a main. Lo único que se interpone es GsDriverEntry, un envoltorio mínimo que el compilador genera automáticamente, hace una pequeña inicialización y llama a tu DriverEntry. Esta es la firma de DriverEntry, en un fragmento de mi guía de fundamentos de drivers:

C
NTSTATUS DriverEntry(
    // tu identidad, te la entregan
    PDRIVER_OBJECT DriverObject,
    // tu clave de configuración en el registro
    PUNICODE_STRING RegistryPath
);

Tres reglas acompañan a esa firma:

  • DriverEntry corre a PASSIVE_LEVEL, el nivel más bajo, el mismo que un hilo normal.
  • Si devuelve un error en lugar de STATUS_SUCCESS, el driver no se carga y su rutina de descarga nunca se llama. Todo lo que hayas creado antes de fallar lo tienes que deshacer tú, allí mismo.
  • El DRIVER_OBJECT no es tuyo: te lo entregan, y nunca lo borras.

IRQL, la ley física del kernel#

El IRQL (Interrupt Request Level) es un número pequeño, y hay uno por núcleo: de 0 a 15 en Windows de 64 bits, que es el único que existe en Windows 11, y de 0 a 31 en el antiguo x86 de 32 bits. La regla es sencilla y no admite excepciones: el código que corre a un IRQL solo puede ser interrumpido por algo de IRQL estrictamente mayor. Esta es la tabla de mi guía, con los valores de 64 bits:

Nivel Valor Qué corre ahí Memoria paginable / esperar
PASSIVE_LEVEL 0 Hilos normales, modo usuario, DriverEntry, AddDevice, Unload Sí / Sí
APC_LEVEL 1 Casi idéntico a PASSIVE_LEVEL; solo bloquea la entrega de APC Sí / casi siempre
DISPATCH_LEVEL 2 DPC, rutinas de cancelación, IoTimer; no hay planificación de hilos No / No (bugcheck)
DIRQL 3–12 Rutinas de interrupción (ISR) de cada dispositivo No / No: lo mínimo y volver

Por encima de los dispositivos quedan niveles reservados al propio sistema, como el reloj (13) o el máximo, HIGH_LEVEL (15). La frontera importante, sin embargo, está entre el 1 y el 2. Por debajo, tu código se comporta como cualquier hilo: puede esperar, dormir y tocar memoria que quizá esté en disco, porque si hace falta el sistema la trae. A partir de DISPATCH_LEVEL no hay planificación de hilos en ese núcleo: nada puede dormirte ni despertarte, y el sistema no puede parar a leer una página del disco. Ahí está la explicación del pantallazo del principio: la misma línea de código que funciona a PASSIVE_LEVEL tumba el sistema a DISPATCH_LEVEL si la memoria que toca resulta estar paginada.

El viaje de un IRP#

Cuando un programa pide algo a un dispositivo, esa petición no llega al driver como una llamada de función. Llega como un IRP (I/O Request Packet): una estructura que representa la petición de E/S y que baja por una pila de drivers. Mi guía lo resume en seis pasos:

  1. La aplicación llama a ReadFile, WriteFile o DeviceIoControl.
  2. El I/O Manager construye el IRP, con un IO_STACK_LOCATION para cada driver de la pila.
  3. Los drivers de filtro, si los hay, lo inspeccionan y lo pasan hacia abajo.
  4. El driver de función recibe el IRP en su rutina de dispatch.
  5. El driver de bus habla con el hardware.
  6. IoCompleteRequest devuelve el IRP hacia arriba, con el resultado.

La regla que no se negocia: cada IRP se completa o se reenvía exactamente una vez. Un IRP olvidado no es una fuga de memoria, es una E/S colgada: un programa esperando para siempre una respuesta que no va a llegar. Y mi guía incluye un aviso muy concreto: si llamas a IoSkipCurrentIrpStackLocation y después pones una rutina de finalización, puedes pisar la del driver que está por encima de ti.

Compruébalo desde modo usuario#

No hace falta escribir un driver para ver el primer paso de ese viaje. Este programa abre el primer disco físico y le pide su geometría al driver de disco de Windows con DeviceIoControl. Abre el disco con acceso 0, que solo permite consultar metadatos: no lee ni escribe nada, y por eso no necesita permisos de administrador. Como la entrada de Ring -2, es código solo para Windows:

ioctl.cpp
// g++ -std=c++20 ioctl.cpp -o ioctl.exe   (Windows)
// Pide al driver de disco su geometría con DeviceIoControl,
// sin permisos de administrador.
#include <windows.h>
#include <winioctl.h>
#include <format>
#include <iostream>

int main() {
    // Acceso 0: solo metadatos; no se lee ni se escribe el disco.
    HANDLE disk = CreateFileW(L"\\\\.\\PhysicalDrive0", 0,
                              FILE_SHARE_READ | FILE_SHARE_WRITE,
                              nullptr, OPEN_EXISTING, 0, nullptr);
    if (disk == INVALID_HANDLE_VALUE) return 1;

    DISK_GEOMETRY_EX geo{};
    DWORD bytes = 0;
    const DWORD code = IOCTL_DISK_GET_DRIVE_GEOMETRY_EX;
    const BOOL ok = DeviceIoControl(disk, code, nullptr, 0,
                                    &geo, sizeof geo, &bytes, nullptr);
    CloseHandle(disk);
    if (!ok) return 1;

    // CTL_CODE(DeviceType, Function, Method, Access), desempaquetado
    std::cout << std::format("IOCTL    0x{:08X}\n", code)
              << std::format("  DeviceType 0x{:X}\n", code >> 16)
              << std::format("  Function   0x{:X}\n", (code >> 2) & 0xFFF)
              << std::format("  Method     {}\n", code & 3)
              << std::format("  Access     {}\n", (code >> 14) & 3)
              << std::format("Disk     {:.1f} GB\n",
                             geo.DiskSize.QuadPart / 1e9)
              << std::format("Sector   {} bytes\n",
                             geo.Geometry.BytesPerSector);
}

Esta es la salida real en mi equipo, un Intel Core i9-14900KF con Windows 11:

salidaoutput
IOCTL    0x000700A0
  DeviceType 0x7
  Function   0x28
  Method     0
  Access     0
Disk     2000.4 GB
Sector   512 bytes

Esa llamada a DeviceIoControl es exactamente el paso 1 del viaje. El I/O Manager construye un IRP de tipo IRP_MJ_DEVICE_CONTROL y lo baja por la pila de drivers de disco, que responde con un disco de 2000,4 GB y sectores de 512 bytes.

Lo más interesante está en las primeras líneas. El código 0x000700A0 no es un número arbitrario: es la macro CTL_CODE empaquetada, y el programa la desempaqueta. DeviceType 0x7 es el tipo disco; Function 0x28 es la operación concreta; Method 0 es METHOD_BUFFERED; y Access 0 es FILE_ANY_ACCESS, que es justo lo que permite pedirla con un handle abierto con acceso 0. Enseguida veremos qué significan esos métodos, porque es lo mismo que define un driver cuando crea sus propios códigos.

La tabla de dispatch#

¿Cómo sabe el driver qué función llamar cuando le llega un IRP? Por su tabla de dispatch. DRIVER_OBJECT.MajorFunction es un array de punteros a función indexado por el tipo de petición, IRP_MJ_*, y DriverEntry la rellena. Otro fragmento de mi guía:

C
NTSTATUS DriverEntry(PDRIVER_OBJECT DriverObject,
                     PUNICODE_STRING RegistryPath) {
    DriverObject->MajorFunction[IRP_MJ_CREATE]         = MyCreateClose;
    DriverObject->MajorFunction[IRP_MJ_CLOSE]          = MyCreateClose;
    DriverObject->MajorFunction[IRP_MJ_DEVICE_CONTROL] = MyDeviceControl;
    DriverObject->DriverUnload                         = MyUnload;
    // ... IoCreateDevice e IoCreateSymbolicLink (más abajo) ...
    return STATUS_SUCCESS;
}

Las entradas que no asignas fallan la petición por defecto. Y mi guía insiste en el orden: la tabla se rellena antes de llamar a IoCreateDevice, para que el dispositivo nunca exista con la tabla a medias.

IOCTL propios#

Cuando un driver necesita operaciones que no encajan en leer o escribir, define sus propios IOCTL con la misma macro que acabamos de desempaquetar:

C
#define IOCTL_MY_DEVICE_DO_THING \
    CTL_CODE(FILE_DEVICE_UNKNOWN, 0x800, METHOD_BUFFERED, FILE_ANY_ACCESS)

El tercer argumento, el método de transferencia, decide cómo viajan los datos entre modo usuario y el driver:

  • METHOD_BUFFERED: el I/O Manager copia los datos a un búfer del sistema. Es el recomendado, y el que usa el driver de disco en nuestro ejemplo.
  • METHOD_IN_DIRECT y METHOD_OUT_DIRECT: los datos se describen con una MDL. Pensados para transferencias grandes.
  • METHOD_NEITHER: el driver recibe los punteros de usuario en bruto y tiene que validarlo todo a mano. Se usa poco, y por algo.
La regla de oro

Nunca te fíes del tamaño ni del contenido de un búfer que viene de modo usuario, sea cual sea el método. Es la misma idea que vimos en las syscalls: el kernel no se fía de nada que venga de Ring 3.

Sincronización: ¿encuentras el bug?#

Un driver atiende peticiones de varios hilos y de varios núcleos a la vez, así que necesita proteger sus datos compartidos. Tiene dos herramientas principales, y el IRQL decide cuál puedes usar:

  • Spinlock: se adquiere a DISPATCH_LEVEL o por debajo, y te sube a DISPATCH_LEVEL mientras lo tienes. Si está ocupado, esperas activamente dando vueltas, así que hay que tenerlo el menor tiempo posible.
  • Mutex: para esperarlo hay que estar como mucho a APC_LEVEL, y en la práctica a PASSIVE_LEVEL. Si está ocupado, el hilo se duerme hasta que se libere.

Con eso, este fragmento de mi guía es un reto. Míralo con calma antes de seguir leyendo:

C
VOID MyDpcRoutine(PKDPC Dpc, PVOID Context, PVOID Arg1, PVOID Arg2) {
    KeAcquireSpinLock(&g_DeviceLock, &oldIrql);
    // Recurso compartido que TAMBIÉN usa un manejador de IOCTL
    KeWaitForSingleObject(&g_ConfigMutex, Executive, KernelMode,
                          FALSE, NULL);
    UpdateSharedConfig();
    KeReleaseMutex(&g_ConfigMutex, FALSE);
    KeReleaseSpinLock(&g_DeviceLock, oldIrql);
}

¿Lo ves? Un DPC ya corre a DISPATCH_LEVEL, y además tiene un spinlock adquirido. Esperar en un objeto del dispatcher, como el mutex, a ese IRQL, con una espera que no sea de tiempo cero (aquí es indefinida), es un error fatal documentado: bugcheck inmediato. Tiene todo el sentido si recuerdas la tabla: a DISPATCH_LEVEL no hay planificación de hilos, así que nada puede dormir a este hilo ni despertarlo cuando el mutex quede libre.

El arreglo depende de quién toca el dato:

  • Si lo tocan el DPC y el manejador del IOCTL, se protege con el mismo spinlock en ambos sitios, y el mutex sobra.
  • Si solo debería tocarse a PASSIVE_LEVEL, ese trabajo se saca del DPC con un work item, que el sistema ejecuta después en un hilo a PASSIVE_LEVEL.

Ciclo de vida: crear y descargar#

Para que un programa pueda hablar con el driver, este tiene que crear un dispositivo. Lo hace con IoCreateDevice, cuyos parámetros clave son el tamaño de la extensión del dispositivo (la memoria propia del driver para ese dispositivo), un nombre del tipo \Device\…, el tipo de dispositivo y FILE_DEVICE_SECURE_OPEN. Después, IoCreateSymbolicLink crea un enlace simbólico para que modo usuario pueda abrirlo con CreateFile.

Con un matiz: la documentación de Microsoft recomienda que los drivers de producción usen IoRegisterDeviceInterface, con una clase de interfaz identificada por un GUID. El enlace simbólico es lo habitual en guías y drivers sencillos.

Al descargar, DriverUnload deshace todo en orden inverso: primero borra el enlace simbólico, después el dispositivo, y libera todo lo que haya reservado. En el kernel no hay proceso que cerrar: una fuga dura hasta el reinicio.

Cómo se trabaja de verdad#

Todo lo anterior explica por qué el desarrollo de drivers tiene su propio método de trabajo. Mi guía de laboratorio lo organiza así:

  • Nunca en la máquina principal. Se trabaja en una máquina virtual de pruebas.
  • Test-signing en esa VM, para poder cargar tus drivers de prueba. Queda bloqueado si la VM tiene Secure Boot activo.
  • Depuración con dos máquinas. Un bugcheck mata también a cualquier depurador que corra en la misma máquina, así que WinDbg se conecta desde fuera: por una tubería serie de la VM o por red con KDNET.
  • Símbolos del servidor público de Microsoft, para que las pilas de llamadas tengan nombres y no solo direcciones.
  • Driver Verifier (verifier.exe), que estresa el driver a propósito para que los fallos de memoria, las fugas y las violaciones de IRQL salgan en el laboratorio y no en producción. Solo en la VM.
  • Volcado de memoria del kernel activado, y el reinicio automático desactivado, para poder leer la pantalla azul y analizarla después.
  • Snapshots. Cuando la VM queda rota, se vuelve al último snapshot en vez de intentar salvarla.
Nota personal

Estas dos guías, la de fundamentos de drivers y la de laboratorio, las escribí para mi propio laboratorio, donde compilo los drivers con Visual Studio Build Tools 2022 y el WDK 10.0.26100. Si tuviera que quedarme con una sola lección de las que más se repiten en ellas, sería esta: los pantallazos frecuentes son una parte normal y esperada del desarrollo de drivers. No son un fracaso si el laboratorio está preparado para ellos.

En el kernel no hay red#

Si hay una idea que une todo lo anterior es que en el kernel no hay red. En modo usuario, un puntero malo cierra tu programa; en Ring 0, cierra el sistema. Por eso las reglas de esta entrada no son cuestión de estilo: respetar el IRQL, completar cada IRP exactamente una vez y validar cada búfer que llega de modo usuario es lo que separa un driver de un pantallazo.

Si quieres ver dónde encaja todo esto en la arquitectura de la CPU, la serie de los anillos empieza en De Ring 3 a Ring 0, baja al hipervisor en Ring -1 y termina en el firmware, en Ring -2.

The most common crash is misunderstood#

If you've ever seen a blue screen, chances are it said IRQL_NOT_LESS_OR_EQUAL. That's bug check 0xA, and its name suggests someone broke an abstract rule about interrupt levels. It's almost never that.

In the vast majority of real cases, there's something much more mundane behind it: a bad pointer, NULL or pointing to memory that has already been freed, or pageable memory touched at an IRQL too high for the page fault to be resolved. At DISPATCH_LEVEL or above, the system can't bring pages in from disk. What in a normal program would be a routine page fault, resolved without you ever noticing, becomes a bug check in the kernel. Its relative, 0xD1 (DRIVER_IRQL_NOT_LESS_OR_EQUAL), goes one step further and points straight at a driver.

That's why, when you debug one of these, the useful question isn't "which IRQL rule did I break?", but this one:

The question to ask

Which pointer did I dereference? Was it valid? And was its memory guaranteed to be present, given the IRQL my code was running at?

In From Ring 3 to Ring 0 we described the boundary between your programs and the kernel from the outside: how it's crossed and why a failure on the other side brings down the whole system. Today we cross that boundary and look at how the code living inside is organized: a Windows driver, from the moment it loads until it serves a request. To understand that question, we first need to understand what a driver is and what IRQL is.

A driver isn't a program#

An executable is passive: the operating system starts it, gives it its own memory space and, if it fails, closes the process and cleans up after it. A driver has none of that. It runs as part of the operating system, in Ring 0, and there's no process boundary protecting the rest of the system from its mistakes.

It doesn't have a main either. Its entry point is DriverEntry, and the Windows I/O Manager starts it when it loads the driver, without the C runtime startup chain that ends up calling main in a normal program. The only thing in between is GsDriverEntry, a minimal wrapper the compiler generates automatically, which does a little initialization and calls your DriverEntry. This is DriverEntry's signature, in a fragment from my driver fundamentals guide:

C
NTSTATUS DriverEntry(
    // your identity, handed to you
    PDRIVER_OBJECT DriverObject,
    // your configuration key in the registry
    PUNICODE_STRING RegistryPath
);

Three rules come with that signature:

  • DriverEntry runs at PASSIVE_LEVEL, the lowest level, the same as a normal thread.
  • If it returns an error instead of STATUS_SUCCESS, the driver isn't loaded and its unload routine is never called. Anything you created before failing is yours to undo, right there.
  • The DRIVER_OBJECT isn't yours: it's handed to you, and you never delete it.

IRQL, the kernel's law of physics#

IRQL (Interrupt Request Level) is a small number, and there's one per core: from 0 to 15 on 64-bit Windows, the only kind Windows 11 comes in, and from 0 to 31 on old 32-bit x86. The rule is simple and has no exceptions: code running at a given IRQL can only be interrupted by something at a strictly higher IRQL. This is the table from my guide, with the 64-bit values:

Level Value What runs there Pageable memory / waiting
PASSIVE_LEVEL 0 Normal threads, user mode, DriverEntry, AddDevice, Unload Yes / Yes
APC_LEVEL 1 Almost identical to PASSIVE_LEVEL; only blocks APC delivery Yes / almost always
DISPATCH_LEVEL 2 DPCs, cancel routines, IoTimer; no thread scheduling No / No (bug check)
DIRQL 3–12 Each device's interrupt service routines (ISRs) No / No: do the minimum and return

Above the devices sit levels reserved for the system itself, such as the clock (13) or the highest one, HIGH_LEVEL (15). The important boundary, though, is between 1 and 2. Below it, your code behaves like any thread: it can wait, sleep and touch memory that might be on disk, because the system brings it in if needed. From DISPATCH_LEVEL up there's no thread scheduling on that core: nothing can put you to sleep or wake you up, and the system can't stop to read a page from disk. That's the explanation for the crash at the start: the same line of code that works at PASSIVE_LEVEL brings the system down at DISPATCH_LEVEL if the memory it touches turns out to be paged out.

An IRP's journey#

When a program asks a device for something, that request doesn't reach the driver as a function call. It arrives as an IRP (I/O Request Packet): a structure that represents the I/O request and travels down a driver stack. My guide sums it up in six steps:

  1. The application calls ReadFile, WriteFile or DeviceIoControl.
  2. The I/O Manager builds the IRP, with one IO_STACK_LOCATION for each driver in the stack.
  3. Filter drivers, if there are any, inspect it and pass it down.
  4. The function driver receives the IRP in its dispatch routine.
  5. The bus driver talks to the hardware.
  6. IoCompleteRequest sends the IRP back up, with the result.

The non-negotiable rule: every IRP is completed or forwarded exactly once. A forgotten IRP isn't a memory leak, it's hung I/O: a program waiting forever for an answer that will never come. And my guide includes a very specific warning: if you call IoSkipCurrentIrpStackLocation and then set a completion routine, you can overwrite the one belonging to the driver above you.

Check it from user mode#

You don't need to write a driver to see the first step of that journey. This program opens the first physical disk and asks the Windows disk driver for its geometry with DeviceIoControl. It opens the disk with access 0, which only allows querying metadata: it doesn't read or write anything, which is why it doesn't need administrator rights. Like the Ring -2 post, this is Windows-only code:

ioctl.cpp
// g++ -std=c++20 ioctl.cpp -o ioctl.exe   (Windows)
// Asks the disk driver for its geometry with DeviceIoControl,
// without administrator rights.
#include <windows.h>
#include <winioctl.h>
#include <format>
#include <iostream>

int main() {
    // Access 0: metadata only; the disk is neither read nor written.
    HANDLE disk = CreateFileW(L"\\\\.\\PhysicalDrive0", 0,
                              FILE_SHARE_READ | FILE_SHARE_WRITE,
                              nullptr, OPEN_EXISTING, 0, nullptr);
    if (disk == INVALID_HANDLE_VALUE) return 1;

    DISK_GEOMETRY_EX geo{};
    DWORD bytes = 0;
    const DWORD code = IOCTL_DISK_GET_DRIVE_GEOMETRY_EX;
    const BOOL ok = DeviceIoControl(disk, code, nullptr, 0,
                                    &geo, sizeof geo, &bytes, nullptr);
    CloseHandle(disk);
    if (!ok) return 1;

    // CTL_CODE(DeviceType, Function, Method, Access), unpacked
    std::cout << std::format("IOCTL    0x{:08X}\n", code)
              << std::format("  DeviceType 0x{:X}\n", code >> 16)
              << std::format("  Function   0x{:X}\n", (code >> 2) & 0xFFF)
              << std::format("  Method     {}\n", code & 3)
              << std::format("  Access     {}\n", (code >> 14) & 3)
              << std::format("Disk     {:.1f} GB\n",
                             geo.DiskSize.QuadPart / 1e9)
              << std::format("Sector   {} bytes\n",
                             geo.Geometry.BytesPerSector);
}

This is the real output on my machine, an Intel Core i9-14900KF running Windows 11:

salidaoutput
IOCTL    0x000700A0
  DeviceType 0x7
  Function   0x28
  Method     0
  Access     0
Disk     2000.4 GB
Sector   512 bytes

That DeviceIoControl call is exactly step 1 of the journey. The I/O Manager builds an IRP_MJ_DEVICE_CONTROL IRP and sends it down the disk driver stack, which answers with a 2000.4 GB disk and 512-byte sectors.

The most interesting part is in the first lines. The code 0x000700A0 isn't an arbitrary number: it's the CTL_CODE macro packed together, and the program unpacks it. DeviceType 0x7 is the disk type; Function 0x28 is the specific operation; Method 0 is METHOD_BUFFERED; and Access 0 is FILE_ANY_ACCESS, which is precisely what lets us request it with a handle opened with access 0. We'll see what those methods mean in a moment, because it's the same thing a driver defines when it creates its own codes.

The dispatch table#

How does the driver know which function to call when an IRP arrives? Through its dispatch table. DRIVER_OBJECT.MajorFunction is an array of function pointers indexed by request type, IRP_MJ_*, and DriverEntry fills it in. Another fragment from my guide:

C
NTSTATUS DriverEntry(PDRIVER_OBJECT DriverObject,
                     PUNICODE_STRING RegistryPath) {
    DriverObject->MajorFunction[IRP_MJ_CREATE]         = MyCreateClose;
    DriverObject->MajorFunction[IRP_MJ_CLOSE]          = MyCreateClose;
    DriverObject->MajorFunction[IRP_MJ_DEVICE_CONTROL] = MyDeviceControl;
    DriverObject->DriverUnload                         = MyUnload;
    // ... IoCreateDevice and IoCreateSymbolicLink (see below) ...
    return STATUS_SUCCESS;
}

Entries you don't assign fail the request by default. And my guide insists on the order: the table is filled in before calling IoCreateDevice, so the device never exists with a half-filled table.

Custom IOCTLs#

When a driver needs operations that don't fit reading or writing, it defines its own IOCTLs with the same macro we just unpacked:

C
#define IOCTL_MY_DEVICE_DO_THING \
    CTL_CODE(FILE_DEVICE_UNKNOWN, 0x800, METHOD_BUFFERED, FILE_ANY_ACCESS)

The third argument, the transfer method, decides how data travels between user mode and the driver:

  • METHOD_BUFFERED: the I/O Manager copies the data into a system buffer. It's the recommended one, and the one the disk driver uses in our example.
  • METHOD_IN_DIRECT and METHOD_OUT_DIRECT: the data is described with an MDL. Meant for large transfers.
  • METHOD_NEITHER: the driver receives raw user pointers and has to validate everything by hand. It's rarely used, and for good reason.
The golden rule

Never trust the size or the contents of a buffer coming from user mode, whatever the method. It's the same idea we saw with syscalls: the kernel trusts nothing that comes from Ring 3.

Synchronization: can you spot the bug?#

A driver serves requests from several threads and several cores at once, so it needs to protect its shared data. It has two main tools, and IRQL decides which one you can use:

  • Spinlock: acquired at DISPATCH_LEVEL or below, and it raises you to DISPATCH_LEVEL while you hold it. If it's taken, you wait actively by spinning, so it has to be held for as short a time as possible.
  • Mutex: waiting on it requires APC_LEVEL at most, and in practice PASSIVE_LEVEL. If it's taken, the thread sleeps until it's released.

With that, this fragment from my guide is a challenge. Take a good look before reading on:

C
VOID MyDpcRoutine(PKDPC Dpc, PVOID Context, PVOID Arg1, PVOID Arg2) {
    KeAcquireSpinLock(&g_DeviceLock, &oldIrql);
    // Shared resource ALSO used by an IOCTL handler
    KeWaitForSingleObject(&g_ConfigMutex, Executive, KernelMode,
                          FALSE, NULL);
    UpdateSharedConfig();
    KeReleaseMutex(&g_ConfigMutex, FALSE);
    KeReleaseSpinLock(&g_DeviceLock, oldIrql);
}

See it? A DPC already runs at DISPATCH_LEVEL, and on top of that it holds a spinlock. Waiting on a dispatcher object, such as the mutex, at that IRQL, with anything but a zero timeout (here it's unbounded), is a documented fatal error: an immediate bug check. It makes perfect sense if you remember the table: at DISPATCH_LEVEL there's no thread scheduling, so nothing can put this thread to sleep or wake it up when the mutex becomes free.

The fix depends on who touches the data:

  • If both the DPC and the IOCTL handler touch it, protect it with the same spinlock in both places, and the mutex goes away.
  • If it should only be touched at PASSIVE_LEVEL, move that work out of the DPC with a work item, which the system runs later on a thread at PASSIVE_LEVEL.

Lifecycle: create and unload#

For a program to talk to the driver, the driver has to create a device. It does so with IoCreateDevice, whose key parameters are the size of the device extension (the driver's own memory for that device), a name of the form \Device\…, the device type and FILE_DEVICE_SECURE_OPEN. Then IoCreateSymbolicLink creates a symbolic link so that user mode can open it with CreateFile.

With one caveat: Microsoft's documentation recommends that production drivers use IoRegisterDeviceInterface, with an interface class identified by a GUID. The symbolic link is the usual choice in guides and simple drivers.

On unload, DriverUnload undoes everything in reverse order: first it deletes the symbolic link, then the device, and it frees everything it allocated. In the kernel there's no process to close: a leak lasts until reboot.

How it's really done#

All of the above explains why driver development has its own way of working. My lab guide organizes it like this:

  • Never on your main machine. You work in a test virtual machine.
  • Test-signing in that VM, so you can load your test drivers. It's blocked if the VM has Secure Boot enabled.
  • Two-machine debugging. A bug check also kills any debugger running on the same machine, so WinDbg connects from outside: over the VM's serial pipe or over the network with KDNET.
  • Symbols from Microsoft's public server, so call stacks show names and not just addresses.
  • Driver Verifier (verifier.exe), which stresses the driver on purpose so memory errors, leaks and IRQL violations show up in the lab and not in production. Only in the VM.
  • Kernel memory dump enabled and automatic restart disabled, so you can read the blue screen and analyze it afterwards.
  • Snapshots. When the VM breaks, you go back to the last snapshot instead of trying to rescue it.
Personal note

I wrote both guides, the driver fundamentals one and the lab one, for my own lab, where I build drivers with Visual Studio Build Tools 2022 and the WDK 10.0.26100. If I had to keep just one of the lessons that come up most often in them, it would be this: frequent crashes are a normal and expected part of driver development. They're not a failure if the lab is ready for them.

There's no safety net in the kernel#

If one idea ties all of the above together, it's that there's no safety net in the kernel. In user mode, a bad pointer closes your program; in Ring 0, it takes down the system. That's why the rules in this post aren't a matter of style: respecting IRQL, completing every IRP exactly once and validating every buffer that comes from user mode is what separates a driver from a blue screen.

If you want to see where all this fits in the CPU architecture, the rings series starts at From Ring 3 to Ring 0, goes down to the hypervisor in Ring -1 and ends in the firmware, in Ring -2.

ReferenciasReferences

  1. Writing a DriverEntry RoutineMicrosoft · learn.microsoft.com
  2. Managing Hardware PrioritiesMicrosoft · learn.microsoft.com
  3. Bug Check 0xA: IRQL_NOT_LESS_OR_EQUALMicrosoft · learn.microsoft.com
  4. Bug Check 0xD1: DRIVER_IRQL_NOT_LESS_OR_EQUALMicrosoft · learn.microsoft.com
  5. I/O Request PacketsMicrosoft · learn.microsoft.com
  6. Defining I/O Control CodesMicrosoft · learn.microsoft.com
  7. Buffer Descriptions for I/O Control CodesMicrosoft · learn.microsoft.com
  8. DeviceIoControl function (ioapiset.h)Microsoft · learn.microsoft.com
  9. IOCTL_DISK_GET_DRIVE_GEOMETRY_EXMicrosoft · learn.microsoft.com
  10. Driver VerifierMicrosoft · learn.microsoft.com

CompartirShare

¿Tu proyecto necesita ir por debajo de la capa de aplicación?Does your project need to go below the application layer?

Software de sistemas, rendimiento y seguridad en Windows y Linux: diagnóstico, herramientas a medida y revisión de código de bajo nivel.Systems software, performance and security on Windows and Linux: diagnostics, custom tooling and low-level code review.

HablemosLet's talk→

ComentariosComments

Para comentar necesitas una cuenta de GitHub. Los comentarios se guardan en GitHub Discussions (SoftDryzz/blog-comments) y se moderan: no se publican enlaces promocionales.Commenting requires a GitHub account. Comments live in GitHub Discussions (SoftDryzz/blog-comments) and are moderated: no promotional links.

Sigue leyendoKeep reading

Bajo nivel · Parte 3Low-level · Part 3

Ring -2: SMI, SMRAM y RSM, cómo el firmware toma la CPU por debajo del hipervisorRing -2: SMI, SMRAM and RSM, how firmware takes over the CPU below the hypervisor

Qué hay por debajo del hipervisor: System Management Mode, el código del firmware que se ejecuta a escondidas del sistema, y cómo comprobar desde C++ qué protecciones declara tu firmware.What lives below the hypervisor: System Management Mode, the firmware code that runs hidden from the OS, and how to check from C++ which protections your firmware declares.

· 12 min

Bajo nivel · Parte 2Low-level · Part 2

Ring -1: el hipervisor que vigila a tu kernelRing -1: the hypervisor watching over your kernel

Qué hay por debajo del kernel: el hipervisor, en modo VMX root (modo host de SVM en AMD). Cómo funciona con VT-x, AMD-V y EPT, por qué Windows usa uno para protegerse a sí mismo y cómo comprobarlo desde C++ con CPUID.What lives below the kernel: the hypervisor, in VMX root mode (SVM host mode on AMD). How it works with VT-x, AMD-V and EPT, why Windows uses one to protect itself, and how to check it from C++ with CPUID.

· 10 min