AI Model InferenceΒΆ

OverviewΒΆ

The AI Bunker reference design flavor supports running AI inference workloads in Bunker OS while keeping Open World as the caller. This feature is available on supported platforms with compatible AI frameworks. The separation is intentional: models, runtime libraries, and hardware accelerators reside in the trusted execution environment of Bunker OS, protected from the less-controlled Open World.

The platform supports multiple AI frameworks with two integration approaches:

RPC-Based Integration (Turmux over vsock)ΒΆ

Most frameworks use the Turmux RPC pattern: a provider process in Bunker OS loads models, registers RPC methods with the Turmux router, and executes inference on behalf of Open World callers.

Open World applications request inference over the vsock channel through Turmux, receive results back through the same channel, and remain completely unaware of the model file path, the runtime version, or any hardware-specific details of the inference engine.

The protocol between caller and provider is Protobuf over Turmux, exactly as described in Reference RPC Implementation: Turmux. Supported frameworks include TensorFlow Lite (TFLite) and the ST AI framework from STMicroelectronics.

A key design goal of the TFLite integration is backward compatibility. Existing Python code that already uses the tf.lite.Interpreter or tf.contrib.lite.Interpreter API can be redirected to Bunker OS inference by changing a single import statement. The rpcai compatibility layer translates the standard TFLite Python API into Turmux RPC calls transparently, so that application code requires no other modification.

Delegate-Level Integration (Direct Runtime Delegation)ΒΆ

Vitis AI (from AMD Xilinx) is also supported with a specialized low-level integration: the XRT (Xilinx Runtime) library is forked and modified to delegate certain operations to Bunker OS while keeping the application unaware of the delegation. This approach uses shared memory segments for 0-copy data exchange, eliminating RPC framing overhead and providing superior efficiency for throughput-intensive workloads. Applications require no code changes, they simply link against the modified XRT library. This level of integration requires substantial engineering effort to maintain in-sync patches across runtime updates. If you are interested in delegate-level support for other frameworks, please contact support.

Implementation DetailsΒΆ

The implementation, protocol definitions, integration examples, and installation instructions are available in the AI Inference Implementation Details.