Writing a WPE platform implementation

WPE WebKit renders web content, but it does not talk to the underlying platform on its own. That job belongs to a platform implementation: a small set of GObject subclasses that connect WebKit to a display server, a compositor, a KMS device — whatever your platform provides. This tutorial walks through what a platform implementation does, concept by concept, roughly in the order you would build one.

The contract is three classes:

  • WPEDisplay — the connection to the platform, and the factory that creates the other objects.
  • WPEToplevel — the surface a view is presented in; on a windowed platform, a native window.
  • WPEView — renders WebKit’s content into a toplevel and receives input for it.

Most platform implementations target a windowed system; Wayland is the common case on Linux. A windowless, offscreen implementation — like the in-tree headless one — is possible too, but it is the exception. Throughout this page each concept is illustrated with whichever built-in implementation shows it most clearly, usually Wayland and occasionally headless, pointing to each one’s full source for the complete code. The out-of-tree GTK implementation is a good example of a real, non-trivial platform built entirely on the public API.

A platform implementation is made available in one of two ways: as a loadable module that WebKit discovers automatically, or linked directly into an application that constructs the WPEDisplay itself. The module is optional — this page focuses on writing the implementation; how WebKit discovers modules at runtime is a separate topic.

The conceptual model behind these classes is introduced in Overview. The examples here use the public wpe-platform-2.0 API, in C.

If you are porting an existing WPEBackend-fdo backend, the “exportable” callback set (export_fdo_egl_image and friends) collapses into the single WPEPlatform.ViewClass.render_buffer below, and the rest of the fdo surface maps onto the classes described here. The Migration mapping table pairs each old symbol with its WPEPlatform equivalent.

The display: connecting and creating objects

The WPEDisplay is the entry point. It opens the connection to the platform, creates WPEView and WPEToplevel instances on request, and optionally exposes extras such as an EGLDisplay, a WPEKeymap, or the set of WPEScreens.

You implement it as a GObject subclass. Declaring and registering GObject types is standard GLib boilerplate — see the GObject documentation if it is unfamiliar — so only the WPE-specific parts are shown here.

Two vfuncs are mandatory. WPEPlatform.DisplayClass.connect opens the connection to the platform:

static gboolean
my_display_connect (WPEDisplay *display, GError **error)
{
    // Connect to the native platform (a compositor, a device, ...).
    // On failure, set error and return FALSE.
    return TRUE;
}

WPEPlatform.DisplayClass.create_view returns a new WPEView tied to the display; implementations create their own subclass:

static WPEView *
my_display_create_view (WPEDisplay *display)
{
    return WPE_VIEW (g_object_new (MY_TYPE_VIEW, "display", display, NULL));
}

You wire these — and any optional vfuncs — into the class in the usual class_init.

Beyond the mandatory pair, WPEDisplayClass declares a number of optional slots, each backing a capability your platform may or may not have:

Only some of the optional slots have a fallback (the keymap, clipboard, gamepad manager and, when built with libdrm, the preferred buffer formats). The others simply report the capability as missing: without get_egl_display there is no EGLDisplay (a WPE_DISPLAY_ERROR_NOT_SUPPORTED error), without create_toplevel views get no toplevel, and without the screen vfuncs the display has no screens.

When a vfunc fails, populate the GError with the WPEDisplayError domain and an appropriate code, such as WPE_DISPLAY_ERROR_NOT_SUPPORTED or WPE_DISPLAY_ERROR_CONNECTION_FAILED:

g_set_error (error, WPE_DISPLAY_ERROR, WPE_DISPLAY_ERROR_CONNECTION_FAILED,
             "Could not connect to the display server");

The toplevel: the surface

A WPEToplevel is the surface a view is presented in — a native window on a windowed platform, where most of the work is: tracking size, reporting state (active, fullscreen, maximized), setting a title, and forwarding every change back to WebKit. A windowless platform still has a toplevel; it just has no real window behind it.

Your toplevel overrides the vfuncs for the capabilities the platform supports — WPEPlatform.ToplevelClass.resize, WPEPlatform.ToplevelClass.set_fullscreen, WPEPlatform.ToplevelClass.set_maximized, WPEPlatform.ToplevelClass.set_title — and leaves the rest unset. On Wayland these map onto xdg-shell requests; a windowless implementation has no real window, so most of them become state-only or no-ops.

Whenever the toplevel’s size or state changes, tell WebKit with wpe_toplevel_resized() and wpe_toplevel_state_changed():

wpe_toplevel_resized (toplevel, width, height);
wpe_toplevel_state_changed (toplevel,
    WPE_TOPLEVEL_STATE_FULLSCREEN | WPE_TOPLEVEL_STATE_ACTIVE);

A toplevel hosts one or more views. Simple platforms allow a single view that always fills the toplevel, so resizing the toplevel maps directly to resizing that one view. Platforms with window chrome — a menu bar, decorations — or several views per window (as the GTK implementation allows) do not have that one-to-one relationship, and must position and size their views themselves — wpe_toplevel_foreach_view() iterates the views a toplevel hosts.

The view: rendering

The WPEView is where WebKit’s content is rendered and where input arrives. Its central vfunc is WPEPlatform.ViewClass.render_buffer: WebKit hands the view a WPEBuffer, the view presents it onto the toplevel’s surface, and then reports back.

static gboolean
my_view_render_buffer (WPEView *view, WPEBuffer *buffer,
                       const WPERectangle *damage_rects, guint n_damage_rects,
                       GError **error)
{
    if (WPE_IS_BUFFER_DMA_BUF (buffer)) {
        EGLImage image = wpe_buffer_import_to_egl_image (buffer, error);
        // ...present the EGLImage on the toplevel's surface...
    } else if (WPE_IS_BUFFER_SHM (buffer)) {
        // ...blit the pixel data from WPE_BUFFER_SHM (buffer)...
    }
    return TRUE;
}

WebKit produces a WPEBufferDMABuf or a WPEBufferSHM (or a WPEBufferAndroid on Android); dispatch on the concrete type. wpe_buffer_import_to_egl_image() turns a DMA-BUF into an EGLImage for hardware-accelerated presentation.

The reporting is a two-step lifecycle, and the distinction between the two steps matters:

  • wpe_view_buffer_rendered() — the buffer has been presented. The frame is now on screen, but the buffer may still be in use (held by the compositor, queued for scanout), so it must not be reused yet.
  • wpe_view_buffer_released() — the buffer is no longer needed, and WebKit may reuse or destroy it.

A typical implementation presents the buffer in render_buffer, calls buffer_rendered once it is committed, and calls buffer_released later, when the platform signals the buffer is free — a Wayland wl_buffer release, a KMS page-flip completing on the next frame, and so on.

On an explicit-sync platform the buffer carries fences instead of blocking. Take its rendering fence (wpe_buffer_take_rendering_fence()) and hand it to whatever consumes the buffer (the in-tree Wayland implementation passes it to the compositor as the acquire fence, and DRM passes it to KMS as the input fence), or wait on it yourself before reading the buffer. If the consumer gives you a release fence (as the Wayland compositor does), attach it with wpe_buffer_set_release_fence() before buffer_released, so WebKit holds off reusing the buffer until your presentation completes. The WPEPlatform.DisplayClass.use_explicit_sync slot advertises the capability.

Visibility, focus, and geometry

A view has two related but distinct notions of visibility:

  • visible (wpe_view_get_visible()) — whether the view is meant to be shown. This can be TRUE even when nothing is on screen, for example while its toplevel is minimized.
  • mapped (wpe_view_get_mapped()) — whether the view is actually being presented right now: visible and not hidden for another reason, such as its toplevel being minimized.

Your implementation drives the mapped state by calling wpe_view_map() and wpe_view_unmap() as those conditions change; WebKit uses it to pause and resume rendering. Report geometry changes with wpe_view_resized(). When one view fills its toplevel, keeping the two in sync just means resizing the view whenever the toplevel resizes.

Report keyboard focus the same way: call wpe_view_focus_in() when the platform gives the view focus and wpe_view_focus_out() when it loses it. The Wayland implementation drives these from the wl_seat keyboard enter/leave events.

Input

Input is central to a windowed platform, and it flows the opposite way from rendering: the implementation receives events from the system and delivers them to the view. You build a WPEEvent with one of the typed constructors — wpe_event_keyboard_new(), wpe_event_pointer_button_new(), wpe_event_pointer_move_new(), wpe_event_scroll_new(), wpe_event_touch_new() — and hand it to the view with wpe_view_event():

g_autoptr(WPEEvent) event =
    wpe_event_keyboard_new (WPE_EVENT_KEYBOARD_KEY_DOWN, view, /* ... */);
wpe_view_event (view, event);

Multi-touch is delivered as one event per touch point, each tagged with a sequence id (wpe_event_touch_get_sequence_id()) so WebKit can track individual points across their lifetime.

Key events are interpreted through a WPEKeymap; if your display provides none, WPEPlatform falls back to an XKB keymap for the pc105 US layout. The Wayland implementation is the reference here — it drives input from the wl_seat family of interfaces and builds its keymap from the compositor. A windowless implementation has no input at all, which is why headless does not implement any of this.

A view can also override WPEPlatform.ViewClass.lock_pointer and WPEPlatform.ViewClass.unlock_pointer to support the Pointer Lock API, and WPEPlatform.ViewClass.set_cursor_from_name to set the cursor shape.

The display's screens

If your platform has monitors, expose them through WPEPlatform.DisplayClass.get_n_screens and WPEPlatform.DisplayClass.get_screen, returning WPEScreen objects. Wayland maps these onto wl_outputs and DRM onto KMS connectors. This is independent of the buffer-format vfuncs above — a platform can have monitors without special format needs, or negotiate formats without having any monitors.

Making the implementation discoverable

WebKit finds implementations through the GIO extension point WPE_DISPLAY_EXTENSION_POINT_NAME: a WPEDisplay subclass registers itself there with a unique name and a priority. The name is what WPE_PLATFORM selects and the priority orders candidates when several are installed: the Wayland implementation uses 0, and the more specialized DRM and headless ones use -100 so they are tried only after Wayland declines.

How you register depends on whether the implementation is compiled in or loaded as a module.

Linked directly. When the type is part of the binary, the G_DEFINE_..._WITH_CODE form registers it against the extension point at type-registration time:

G_DEFINE_FINAL_TYPE_WITH_CODE (MyDisplay, my_display, WPE_TYPE_DISPLAY,
    g_io_extension_point_implement (WPE_DISPLAY_EXTENSION_POINT_NAME,
        g_define_type_id, "myplatform", 0))

An application that links the library can then construct MyDisplay itself, or let wpe_display_get_default() find it.

As a loadable module. A module’s types are tied to the module’s lifetime, so the display must be a dynamic type — G_DEFINE_DYNAMIC_TYPE_EXTENDED, which generates a my_display_register_type() — registered from the module’s load entry point rather than with WITH_CODE:

G_MODULE_EXPORT void
g_io_module_load (GIOModule *module)
{
    my_display_register_type (G_TYPE_MODULE (module));
    g_io_extension_point_implement (WPE_DISPLAY_EXTENSION_POINT_NAME,
        MY_TYPE_DISPLAY, "myplatform", 0);
}

G_MODULE_EXPORT void
g_io_module_unload (GIOModule *module)
{
}

WebKit scans its module directory eagerly, so a g_io_module_query entry point is not needed. GLib also supports module-name-scoped entry points (g_io_<module-name>_load / _unload), which the out-of-tree GTK implementation uses so several modules can share a binary.

Building and installing

Build the implementation against the wpe-platform-2.0 pkg-config module. A loadable module is a shared library installed into the directory WebKit scans:

${LIBDIR}/wpe-platform-2.0/modules/

The exact path is available from the pkg-config module:

pkg-config --variable=moduledir wpe-platform-2.0

Once installed, force WebKit to use it by setting WPE_PLATFORM to the name you registered:

WPE_PLATFORM=myplatform MiniBrowser https://webkit.org

While iterating on a module that is not installed yet, point WebKit at the build directory with WPE_PLATFORMS_PATH.

A minimum viable implementation

The smallest implementation that renders anything is a WPEDisplay that overrides WPEPlatform.DisplayClass.connect, WPEPlatform.DisplayClass.create_view, and WPEPlatform.DisplayClass.create_toplevel, plus a WPEView whose WPEPlatform.ViewClass.render_buffer calls wpe_view_buffer_rendered() and wpe_view_buffer_released() (WebKit stops producing frames otherwise). The view must also be given a size and be mapped: a new view is 0×0 and unmapped, and attaching a toplevel does not change that. A toplevel gets a default size from the settings, so the headless implementation connects to notify::toplevel and calls wpe_view_resized() with the toplevel’s size followed by wpe_view_map(). The headless implementation is close to this minimum, and additionally provides get_egl_display and get_drm_device so that DMA-BUF buffers can be used. From there, add input, screens, and buffer-format negotiation as the platform you are targeting requires.

For complete, working code, read the in-tree implementations under Source/WebKit/WPEPlatform/wpe/wayland/ for the full windowed case, headless/ for the minimal one — and the out-of-tree GTK implementation for a real-world example built on the public API.