Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
48 commits
Select commit Hold shift + click to select a range
df83bf6
docs: design Phi-4 GGUF AIE4 integration
chizamd Sep 11, 2026
4948126
docs: plan Phi-4 GGUF AIE4 integration
chizamd Sep 11, 2026
d1b0a05
build: add optional dynamic corelib 0.3.0 runtime
chizamd Sep 11, 2026
1f67adb
test: strengthen corelib integration guards
chizamd Sep 11, 2026
f2c1dba
feat: add validated Phi-4 Q8_0 GGUF reader
chizamd Sep 11, 2026
5e6423b
fix: harden Phi-4 GGUF validation
chizamd Sep 11, 2026
804b856
test: assert exact Phi-4 mismatch diagnostics
chizamd Sep 11, 2026
0c3c39f
feat: add corelib-backed Phi-4 AIE4 engine
chizamd Sep 11, 2026
ee66551
fix: harden Phi-4 corelib engine state
chizamd Sep 11, 2026
5d50361
feat: route Phi-4 GGUF models through AIE4
chizamd Sep 11, 2026
8f831fe
fix: harden Phi-4 request lifecycle
chizamd Sep 11, 2026
17d7525
fix: preserve legacy Phi-4 tokenizer semantics
chizamd Sep 11, 2026
1319e07
feat: pull Phi-4 GGUF and tokenizer from pinned sources
chizamd Sep 11, 2026
32f7698
fix: preserve Phi-4 downloader compatibility semantics
chizamd Sep 11, 2026
dfdd1f4
test: validate Phi-4 GGUF AIE4 integration
chizamd Sep 11, 2026
1eb82e4
test: strengthen Phi-4 AIE4 integration coverage
chizamd Sep 11, 2026
14f6161
test: exercise production Phi-4 integration seams
chizamd Sep 11, 2026
a6ce095
fix: release idle NPU queue immediately
chizamd Sep 11, 2026
867df75
fixup! build: add optional dynamic corelib 0.3.0 runtime
chizamd Sep 11, 2026
b0a13ec
fixup! feat: add validated Phi-4 Q8_0 GGUF reader
chizamd Sep 11, 2026
a2f8550
fixup! feat: add corelib-backed Phi-4 AIE4 engine
chizamd Sep 11, 2026
ee78f6f
fixup! feat: route Phi-4 GGUF models through AIE4
chizamd Sep 11, 2026
f7eb0c1
fixup! feat: add corelib-backed Phi-4 AIE4 engine
chizamd Sep 11, 2026
cd2223a
fixup! build: add optional dynamic corelib 0.3.0 runtime
chizamd Sep 11, 2026
7347f5b
fixup! feat: route Phi-4 GGUF models through AIE4
chizamd Sep 11, 2026
edd5cad
fixup! feat: add corelib-backed Phi-4 AIE4 engine
chizamd Sep 11, 2026
465f49b
docs: drop the superpowers plan and design pages from the branch
chizamd Sep 11, 2026
c6827a4
fixup! feat: add corelib-backed Phi-4 AIE4 engine
chizamd Sep 12, 2026
87b9ce8
fixup! feat: route Phi-4 GGUF models through AIE4
chizamd Sep 12, 2026
e6fc279
test: add the real-AIE4 acceptance runner
chizamd Sep 12, 2026
8772108
test: stop the acceptance runner mangling its own curl arguments
chizamd Sep 12, 2026
4d39df1
docs: document developer setup and hardware results
chizamd Sep 12, 2026
a2b730a
feat: report model load time, and account for it on the AIE4 path
chizamd Sep 12, 2026
00e7a60
fix: report load time on the server path too, not only the CLI
chizamd Sep 12, 2026
42d148e
docs: account for the 45 s startup, measured rather than assumed
chizamd Sep 12, 2026
372a039
perf: stop re-hashing the model at startup, and give the packer a thr…
chizamd Sep 12, 2026
87e5fc8
docs: startup is 5 s, not 45 s
chizamd Sep 12, 2026
c52cfec
refactor: restructure the srcs
alfxu-amd Sep 15, 2026
fdb9420
feat: platform filter on model_tag
alfxu-amd Sep 15, 2026
fee9d1f
refactor: split phi4 headers by engine
alfxu-amd Sep 15, 2026
2c7e8fa
refactor: parse tokenizer_config.json once and pass it down to the ba…
chizamd Sep 16, 2026
9592fb9
fix: report truncation as truncation, and name a poisoned model once
chizamd Sep 16, 2026
f21c3ca
perf: pack the 161 weights concurrently instead of one at a time
chizamd Sep 16, 2026
c095e51
docs: separate the two machines' numbers
chizamd Sep 16, 2026
18baf93
docs: the backend guide said not to parallelise the creates
chizamd Sep 16, 2026
424edda
feat: cache the packed weights so requantization is paid once per GGUF
chizamd Sep 16, 2026
d41c2fa
feat: reclaim a weight cache that is not going to be used
chizamd Sep 16, 2026
b5a05ac
refactor: name the backends after the silicon they run on
alfxu-amd Sep 17, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
115 changes: 115 additions & 0 deletions docs/docs/benchmarks/phi4_results.md
Original file line number Diff line number Diff line change
Expand Up @@ -41,3 +41,118 @@ AMD Ryzen™ AI 7 350 (Kraken Point) with 32 GB DRAM; performance is comparable
| **Model** | **HW** | **1k** | **2k** | **4k** | **8k** | **16k** | **32k** |
|------------------|--------------------|--------:|--------:|--------:|--------:|---------:|---------:|
| **Phi-4-mini-instruct** | NPU (FLM) | 643 | 787 | 857 | 809 | 644 | 447 |

---

## 🧪 Phi-4-mini-instruct Q8_0 GGUF on AIE4 (`phi4-mini-it:4b`, resolved for AIE4)

These are **descriptive measurements from a single acceptance run**, not a benchmark sweep and not a pass threshold. They are not comparable to the tables above: the prompts here are 4–10 tokens, whereas those tables sweep 1k–32k, so the per-token rates are dominated by fixed overhead rather than by context length.

## Machine A

Measured at the commit named below, which predates the backend restructure and
the concurrent packer. Its load figures therefore describe the **serial**
packer; machine B carries the current ones. The generation figures here were
not re-measured after those changes.

### Provenance

| | |
|---|---|
| Machine | AIE4 development machine A |
| CPU | AMD Ryzen AI engineering sample |
| NPU | `AMD XDNA(TM) NPU` |
| OS | Microsoft Windows 11 Enterprise 10.0.26100 build 26100 |
| Windows power scheme | Balanced (`381b4222-f694-41f0-9685-ff5bb260df2e`). The NPU power mode is separately set to `performance` by FLM at startup. |
| FastFlowLM commit | `87721089097396579ec4529f50616a6c0e1c7b74` |
| corelib commit / ABI | `3c35aebdefa3f0c2255668bab1be5648ece320f8` / `0.3.0` |
| corelib DLL SHA-256 | `f404da219a3cc84d3334c265e09ba7987f0c4bcc1b1cedeac7c3c45a7be2c9ae` |
| GGUF revision | `78eb92a46fc37e6b524df991ed9aca9bc6aa7b80` |
| Tokenizer/config revision | `cfbefacb99257ffa30c83adab238a50856ac3083` |
| Run | 2026-09-12 00:34:19 → 00:52:02, `passed: true`, 0 failures |

### Measurements

| Metric | Value | Conditions |
|---|---|---|
| Model load to serving | **5.1 / 5.3 s** | fresh `flm serve` processes, timed from launch to the first successful `/api/version`. Was 44–49 s at the accepted commit; see below. |
| Cold TTFT | **4.21 s** | first prompt in a fresh process; includes one-time kernel and ELF setup |
| Warm TTFT | **65.0 ms** | subsequent prompts in the same process |
| Decode, REST | **21.3 tok/s** | `/api/chat`, 16 generated tokens |
| Decode, warm CLI session | **35.8 tok/s** | 10 prompts in one loaded process |

### Startup: 45 s → 5 s

The acceptance run measured 44–49 s to serving. Profiling it with `FLM_AIE4_PROFILE_LOAD=1` found two independent costs, both since fixed:

| Phase | Before | After |
|---|---|---|
| Startup integrity check — SHA-256 over the 4 GB GGUF | ~28 s (62%) | **0 s** — not run |
| Weight requantization — 161 objects from Q8_0 | 15.3 / 14.9 s | **2.5 / 3.0 s** |
| Shape plan | 0.05 s | 0.05 s |
| GGUF resolve, host prep, device tensors | < 0.2 s | < 0.2 s |
| **Process launch to serving** | **44.9 / 46.3 s** | **5.1 / 5.3 s** |

The integrity check was re-hashing every pinned file on every launch — a pull-time concern on the startup path. `flm pull` and `flm check` still verify in full; only the run and serve paths were changed to ask for status alone.

The packer was being given a threads hint of 0, which corelib treats as ONE deliberately, so a single create packed on a single thread. Raising the hint brought requantization to 2.5–3.0 s here, within range of the 2.2 s `python/phi4_driver.py` reports for the same 161 weights. The packer has since moved to concurrent creates instead; machine B carries those figures.

Output was re-verified after the change: `2+2` → `4`, `capital of France` → `Paris`, `primary color` → `Red.`, and a correct one-sentence description of AMD.

Separately, and **not** fixed: `calculate_file_sha256` uses a portable pure-C++ SHA-256 with no hardware acceleration, and takes ~28 s over 4 GB where `Get-FileHash` on the same machine takes **3.67 s**. That ~8× gap is not specific to this model or backend — it is still paid by `flm pull` and `flm check` for every model.

**Do not read the per-process cold cycles as throughput.** Ten fresh-process cycles generating 8 tokens each reported 3.70–20.26 tok/s decode and 1.09–3.65 tok/s prefill. Every one of those pays the one-time setup inside its own measurement window, so the average describes start-up cost, not steady-state speed.

The **5.4×** spread between warm TTFT (65 ms) and cold TTFT (4.21 s), and the **1.7×** spread between the REST and warm-CLI decode figures, are both unexplained by anything measured here. Treat single-run differences below roughly 2× as noise.

### Functional results

All from the same run:

- `flm pull` / `flm check` — four pinned files, all SHA-256 verified; the model directory contains exactly those four.
- CLI — 10/10 fresh-process load-and-generate cycles exited 0; `Backend: corelib_aie4_gguf` and the loaded DLL path reported in every one.
- REST — `/api/chat` and `/v1/chat/completions` both 200, streaming and non-streaming.
- Cancellation — an in-flight stream cancelled cleanly; the next request returned 200 on the same server.
- Capacity boundary — a request totalling 4096 tokens is rejected with **HTTP 400** before submission (`rendered prompt has 4 tokens and requested output has 4092 tokens`); a 4095-token request is admitted.
- No CPU or NPU2 fallback appears in the server log at any point.

### Known issue

One `/api/chat` reply to `What is 2+2?` came back as a truncated markdown image URL (`![](https://media.giphy.com/media/kZl76FZgu`, `done_reason: length`) instead of an answer. The identical prompt answered correctly on three other occasions in the same session, including the recovery request in the same run, so this looks like sampling nondeterminism rather than a routing fault — but it is a single-observation defect, it is not understood, and it is recorded rather than smoothed over.

---

## Machine B

Measured with the concurrent packer, after the backend restructure. **Load only**: TTFT, decode throughput and the acceptance matrix have not been re-run on this machine, so machine A remains the only source for those.

### Provenance

| | |
|---|---|
| Machine | AIE4 development machine B |
| NPU | architecture `aie4` |
| CPU | AMD Ryzen AI engineering sample, 20 cores |
| Memory | 32 GB |
| OS | Windows 11 Enterprise 10.0.26100.4652 |
| XRT / NPU driver / NPU firmware | 2.25.0 / 32.0.20214.4161 / 2.6.1.219 |
| corelib | `3c35aebd`, ABI 0.3.0, built on this machine |

### Model load

Weight requantization is effectively the whole of load; everything else — shape planning, GGUF resolution, host preparation, device allocation — stays under a quarter of a second combined.

| Packer | Requantization | Notes |
|---|---|---|
| Serial, one create at a time | **~30 s** | 29.99 / 30.19 / 25.14 s across runs, consistently slow |
| Concurrent, 8 creates in flight | **3.50 – 13.51 s**, median 8.76 s over 6 runs | one interactive run measured 7.37 s |

Two things are worth separating here.

The **consistent** 30 s came from the packer running effectively single-threaded in that session while the same binary was several times faster elsewhere. Taking the parallelism as threads FastFlowLM owns, rather than as a hint passed to the packer, removes that dependence on the surrounding environment.

What remains is **variance, not a fixed cost**: 3.50–13.51 s on an otherwise idle machine, a ~4× spread, and a separate ten-load run saw 6.74–19.41 s. Single measurements of this phase are not meaningful; quote a range. The variance is not explained by anything measured here.

### Not measured here

Load is where this machine was exercised. Cold and warm TTFT, decode throughput, the REST and cancellation matrix, and the capacity boundary were all measured on machine A at an earlier commit and have **not** been reconfirmed here.
34 changes: 34 additions & 0 deletions docs/docs/instructions/cli.md
Original file line number Diff line number Diff line change
Expand Up @@ -232,6 +232,40 @@ flm serve llama3.2:1b --ctx-len 8192

---

### 🔀 Choose an Execution Backend

**A backend is a piece of hardware.** There are two:

| id | hardware | engine |
|---|---|---|
| `aie2p` | Strix / Krackan Point | FastFlowLM's own NPU kernels |
| `aie4` | AIE4 | AMD's ryzenai-corelib |

You normally never set this. A build targets one NPU generation — `FLM_ENABLE_AIE4` selects AIE4, otherwise AIE2P — and `flm run` prints it as `NPU platform: aie2p`. That platform *is* the backend. A model family has at most one engine per generation, so there is nothing to choose between.

The flag exists for overriding the detection, and for the targets that will join this list later:

```shell
flm run phi4-mini-it:4b --backend aie2p
flm serve phi4-mini-it:4b --backend aie4
```

**Precedence**, highest first:

| # | Source | |
|---|---|---|
| 1 | `--backend <id>` | the flag above, or a `"backend"` field on an `/api/chat` or `/api/generate` request |
| 2 | `FLM_BACKEND=<id>` | environment variable, for a whole shell session |
| 3 | the detected NPU | what `flm validate` reports |

A per-request `"backend"` overrides `--backend` for that request, and reloads the model if it differs from the one already loaded — exactly as asking for a different model does.

Asking for a backend this build does not ship fails at load with a message listing what is actually available. There is no silent fallback to another engine or to the CPU. Note that a release is built for one NPU generation: on hardware it was not built for, `run`, `serve` and `bench` refuse up front, and `flm validate` still reports what it found.

> A build carries the engines for one generation only, so there is no flag that moves it to the other catalog entry — that is a different build of FLM.

---

### 🖧 Set Server Port at Launch

Set a custom port at launch:
Expand Down
65 changes: 65 additions & 0 deletions docs/docs/models/phi.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,4 +22,69 @@ parent: Models
flm run phi4-mini-it:4b
```

---

## 🧪 Model Card: Phi-4-mini-instruct on AIE4 (developer preview)

- **Tag:** `phi4-mini-it:4b` — the same tag as the NPU2 build. A build targets one NPU generation (`FLM_ENABLE_AIE4` selects AIE4, otherwise AIE2P / Strix / Krackan Point), and the tag resolves to the artifacts that generation can run. There is no separate AIE4 tag; `flm list` on an AIE4 build shows only the models it can run.
- **Backend:** `aie4` — the backend is the hardware, and on AIE4 FastFlowLM drives it through AMD's `ryzenai_corelib`
- **Source format:** GGUF, read directly. No ONNX model, no tensor manifest, and no converted or packed weight file is produced or shipped.
- **Quantization:** GGML `Q8_0` in the file, requantized to **group 64** while the weights are packed for the device, through corelib's explicit `*_create_gguf_requantized` entry points. This is a **lossy** second quantization step and it is not reversible; output will differ from the Q8_0 source.
- **Usable generation window:** 4095 tokens — the rendered prompt plus the requested output together, so the largest admissible prompt is 4094. An over-capacity request is rejected with HTTP 400 *before* any work is submitted to the device. Note this is far below the model's 128k context; see below for why.
- **Availability:** Windows only, and this is a **developer build**. The AIE4 runtime is not packaged by the MSI or Inno installer; you build against corelib yourself.

On AIE4 this tag pulls from two pinned repositories, because the GGUF publisher does not ship the tokenizer files FastFlowLM's tokenizer frontend consumes:

| File | Repository | Revision |
|---|---|---|
| `Phi-4-mini-instruct.Q8_0.gguf` | [`unsloth/Phi-4-mini-instruct-GGUF`](https://huggingface.co/unsloth/Phi-4-mini-instruct-GGUF) | `78eb92a46fc37e6b524df991ed9aca9bc6aa7b80` |
| `tokenizer.json` | [`microsoft/Phi-4-mini-instruct`](https://huggingface.co/microsoft/Phi-4-mini-instruct) | `cfbefacb99257ffa30c83adab238a50856ac3083` |
| `tokenizer_config.json` | [`microsoft/Phi-4-mini-instruct`](https://huggingface.co/microsoft/Phi-4-mini-instruct) | `cfbefacb99257ffa30c83adab238a50856ac3083` |
| `config.json` | [`microsoft/Phi-4-mini-instruct`](https://huggingface.co/microsoft/Phi-4-mini-instruct) | `cfbefacb99257ffa30c83adab238a50856ac3083` |

All four are SHA-256 verified before the download is promoted, and the pulled directory contains exactly these four files.

### Building

The AIE4 path is compiled only when you ask for it. With the option off, the binary contains no reference to corelib at all.

From `FastFlowLM/src`, in a Visual Studio developer command prompt:

```powershell
$env:RYZENAI_CORELIB_INCLUDE_DIR = 'C:/path/to/ryzenai-corelib/install/include'
cmake --preset windows-aie4 # sets FLM_ENABLE_AIE4=ON, builds into src/build-aie4
cmake --build --preset windows-aie4
```

The `windows-aie4` preset reads `RYZENAI_CORELIB_INCLUDE_DIR` and `RYZENAI_CORELIB_LIBRARY` from the environment, so set both before configuring. The configure step also locates a Boost include directory, and hard-errors if the option is enabled on a non-Windows host. Everything else — XRT, FFmpeg, curl, FFTW — is the ordinary FastFlowLM dependency set; the AIE4 option does not relax any of it.

### Pointing FastFlowLM at the runtime

An AIE4 build (`-DFLM_ENABLE_AIE4=ON`) **links corelib in**, because the NPU device the whole process shares comes from corelib's `ryzenai::corelib::GetDevice()` rather than from a device FastFlowLM opens itself. Point the build at the library with `RYZENAI_CORELIB_INCLUDE_DIR` and `RYZENAI_CORELIB_LIBRARY`. `FLM_AIE4_CORELIB_PATH` selects a DLL only in the older dynamically loading configuration; in an AIE4 build it is ignored, and `flm` says so if it is set.

The corelib ABI is still pre-1.0, so FastFlowLM requires an **exact `0.3.0`** match on major, minor and patch. The version is queried before any other entry point, so a mismatched runtime reports a version error rather than a missing symbol. Corelib's own dependency directory must be reachable on `PATH`.

```powershell
flm pull phi4-mini-it:4b
flm run phi4-mini-it:4b
```

### Why the context is 4096, and why the usable window is one less

Phi-4-mini itself supports 128k, and the existing `phi4-mini-it:4b` tag defaults to 32k. This backend gives you 4095. That is a real functional regression and it has two separate causes, which are worth keeping apart.

**The 4096 ceiling is a correctness boundary, not a buffer size.** 4096 is exactly Phi-4-mini's `rope.scaling.original_context_length`. LongRoPE selects its factors by *sequence length*, not per position: at or below the original length the short factors apply, above it the long ones do. This implementation derives only the short branch, so 4096 is the point past which the rope tables would silently be wrong. It is enforced rather than assumed — loading fails with `invalid Phi-4 RoPE metadata` unless the GGUF reports `rope.scaling.original_context_length` of exactly 4096. Raising this ceiling means deriving the long factors, not enlarging an array.

**The extra −1 is this frontend's own conservatism.** `kMaxDecodeWindow` is 4095, one below the attention window, so that any request the server admits is guaranteed to have room to finish rather than failing partway. It costs exactly one token and it is not imposed by corelib.

### No fallback

Backend selection follows the build, never a filename or a quantization level. The generation this binary was built for decides two things at once: *which catalog entry* the tag resolves to — the NPU2/Q4NX entry on aie2p, this one on aie4 — and which backend runs it, because the backend id and the platform id are the same string. Once this entry is selected, there is no fallback: if corelib is missing, unloadable, or the wrong version, the tag **fails to load with a diagnostic** rather than quietly running on CPU or on the NPU2/Q4NX backend.

### Naming the backend yourself

The two engines are registered under the hardware they run on, `aie2p` and `aie4`, and you can name one with `--backend`, with `FLM_BACKEND`, or with a `"backend"` field on an `/api/chat` or `/api/generate` request. The [CLI reference](../instructions/cli.md) has the full precedence table.

This does not widen what the hardware can run. A given release is built for one NPU generation, so only that generation's backend is compiled in; asking for the other one fails immediately, naming what this build actually has, instead of failing deep inside an engine that was never going to work. If what you meant was the other catalog entry, that is a different build of FLM, not a different flag.

---
42 changes: 42 additions & 0 deletions src/CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -48,6 +48,27 @@ set(CMAKE_RUNTIME_OUTPUT_DIRECTORY_RELEASE ${CMAKE_RUNTIME_OUTPUT_DIRECTORY})
# ———————————————————————————————————————————————
option(FLM_USE_HRX "Use the HRX amdxdna NPU runtime instead of XRT (0=XRT default, 1=HRX)" OFF)
option(FLM_PORTABLE_BUILD "Build portable distribution with bundled runtime libraries" OFF)
# The two NPU generations do not share an engine: aie2p (Strix / Krackan) runs
# FastFlowLM's own kernels, aie4 runs ryzenai-corelib. A build carries
# one of the two, so this option selects the target hardware, not an extra.
option(FLM_ENABLE_AIE4
"Target AIE4 through ryzenai-corelib instead of aie2p" OFF)

if(FLM_ENABLE_AIE4)
if(NOT WIN32)
message(FATAL_ERROR "FLM_ENABLE_AIE4 currently requires Windows")
endif()
find_path(RYZENAI_CORELIB_INCLUDE_DIR NAMES ryzenai/corelib.h REQUIRED)
# The main binary needs ryzenai::corelib::GetDevice(), a C++ entry point the
# C ABI does not expose, so corelib is linked in rather than dlopened.
find_library(RYZENAI_CORELIB_LIBRARY NAMES ryzenai_corelib corelib
HINTS "${RYZENAI_CORELIB_INCLUDE_DIR}/../lib"
"$ENV{RYZENAI_CORELIB_LIB_DIR}" REQUIRED)
find_path(FLM_CORELIB_BOOST_INCLUDE_DIR NAMES boost/any.hpp
HINTS "$ENV{CONDA_PREFIX}/Library/include"
"$ENV{USERPROFILE}/anaconda3/Library/include"
"C:/dev/boost_1_88_0" REQUIRED)
endif()

if(FLM_USE_HRX)
set(FLM_RUNTIME_NAME "hrx")
Expand Down Expand Up @@ -237,6 +258,12 @@ add_subdirectory(${CMAKE_SOURCE_DIR}/../third_party/tokenizers-cpp
# ———————————————————————————————————————————————
file(GLOB SOURCES "src/*.cpp" "runner/*.cpp" "common/*.cpp" "common/*/*.cpp" "server/*.cpp" "pull/*.cpp" )
file(GLOB HEADERS "include/*.hpp" "runner/*.hpp" "common/*.hpp" "common/*/*.hpp" "server/*.hpp" "pull/*.hpp")
list(FILTER SOURCES EXCLUDE REGEX ".*/common/aie4/.*\\.cpp$")

# Model sources live two levels deeper than the globs above reach, so pull
# them in explicitly; the aie4 half is built separately, below.
include("${CMAKE_SOURCE_DIR}/common/models/models_sources.cmake")
list(APPEND SOURCES ${FLM_MODELS_AIE2P_SOURCES})

# Exclude files that depend on missing libraries for Linux
if(NOT WIN32)
Expand Down Expand Up @@ -269,6 +296,21 @@ endif()

add_executable(flm ${SOURCES} ${HEADERS})

if(FLM_ENABLE_AIE4)
include("${CMAKE_SOURCE_DIR}/common/aie4/aie4_sources.cmake")
add_library(flm_aie4 STATIC ${FLM_AIE4_SOURCES})
target_include_directories(flm_aie4 PUBLIC
"${CMAKE_SOURCE_DIR}/include" "${RYZENAI_CORELIB_INCLUDE_DIR}"
"${XRT_INCLUDE_DIR}" "${FLM_CORELIB_BOOST_INCLUDE_DIR}")
# RYZENAI_CORELIB_STATIC drops the vendor header's dllimport decoration;
# FLM_CORELIB_LINK_STATIC is what tells our own code the symbols are linked
# in and must be bound directly instead of through LoadLibrary.
target_compile_definitions(flm_aie4 PUBLIC
FLM_ENABLE_AIE4=1 RYZENAI_CORELIB_STATIC=1 FLM_CORELIB_LINK_STATIC=1)
target_link_libraries(flm_aie4 PUBLIC "${RYZENAI_CORELIB_LIBRARY}")
target_link_libraries(flm PRIVATE flm_aie4)
endif()

if(WIN32)
if(VCPKG_TOOLCHAIN)
# A vcpkg toolchain is active (e.g. the rocm-npu-staging dev.py build or
Expand Down
Loading