Instructions to use MarcinEU/finegrain-box-segmenter-ONNX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers.js
How to use MarcinEU/finegrain-box-segmenter-ONNX with Transformers.js:
// npm i @huggingface/transformers import { pipeline } from '@huggingface/transformers'; // Allocate pipeline const pipe = await pipeline('mask-generation', 'MarcinEU/finegrain-box-segmenter-ONNX');
- finegrain-box-segmenter β ONNX
finegrain-box-segmenter β ONNX
FP32 ONNX export of finegrain/finegrain-box-segmenter (an MVANet high-resolution background remover), ready to run in the browser, Node.js and Python with ONNX Runtime (CPU, WebGPU (recommended), DirectML).
- Base model:
finegrain/finegrain-box-segmenter(MIT) β MVANet, SafeTensors, arXiv:2404.07445 - Architecture: MVANet with a Swin-B backbone, ~94.6 M parameters; static 1024Γ1024 input, batch fixed at 1.
- What this repo adds: the same network exported to ONNX (full-precision floating-point fp32), so it runs without PyTorch/refiners, plus loss-free graph rewrites that make it load in browsers and cut its peak memory by 40-70% (same outputs - see Graph rewrites).
- Why fp32 in ONNX: maximum compatibility with zero precision risk β the fp32 graph runs on CPU, WebGPU and DirectML, and stays bit-identical to the original release; at ~94.6 M parameters it is still cheap enough to run on a local machine.
- Output is bit-identical to the PyTorch reference on CPU (mask MAE
0.000 / 255; random-input logitsmax|Ξ| = 1.7e-5).
Examples
These tests are intentionally made to be difficult for the models, doing their best to expose the weak points of each model.
- Sub-pixel hair strands
- Refraction + small object touching frame edge
- Large shape touching edge in a style rarely seen in the training data (impasto), no obvious dominant subject
- 3D/miniature with glows and object blending under water
- Cartoon / flat 2D
- Complex fire pattern
- Reflections on a confusing background
- Intricate many-holed topology
The short version: finegrain-box-segmenter is strongest on soft, semi-transparent edges β it keeps more sub-pixel hair strands (1), cleaner glass/refraction edges (2) and more of the flame and fur tips (6) than the other tools, and it resolves the many-holed branch topology (8) cleanly. Its weak spots in this set: the ultra-thin rigging wires in (4) (the BiRefNet variants preserve them better) and the unusual impasto style of (3), where part of the rooster's comb goes semi-transparent. All cutouts were produced the same way: each tool's standard whole-image pipeline β no boxes or prompts β on identical source images, and every crop is the same pixel window. The competitors (finegrain is this model): ben2 = BEN2, bria = BRIA RMBG-2.0, birefnet = BiRefNet, the massive-trained checkpoint (BiRefNet-massive-TR_DIS5K_TR_TEs-epoch_420, run via rembg).
| Original | finegrain | ben2 | bria | birefnet | |
|---|---|---|---|---|---|
| 1 | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | |
| 2 | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | |
| 3 | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | |
| 4 | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | |
| 5 | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | |
| 6 | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | |
| 7 | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | |
| 8 | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() |
Comparative numbers
TL;DR: on CPU the ONNX runs about 20% faster than the PyTorch original (~8 vs ~10 s per image in Node) with about a third of its RAM (1.4 vs 3.8 GB). For a maximum performance use the Node version - it wins both in CPU and (for dedicated GPUs) DirectML, using much less memory and less or about the same time as the original Safetensors on PyTorch (this is mostly due to faster pre- and post-processing). 8k image can be cut out in less than half a second on DirectML through onnxruntime-node 1.26.0 using the reference hardware. PHP version shows the worst result across backends, but that is mostly due to the GD library's limitations rather than the inference itself, so it should be seriously considered as a feasible alternative given more performant image processing. Between the web browsers Chrome offers the best performance, with that not even being close between the two, but (of course) I highly recommend you to avoid ever using this model client-side, as it consumes amount of memory none of the users would expect and risk out-of-memory errors or crashes on mobiles.
| Metric | fp8 Safetensors | fp32 ONNX (this repo) | ||||
|---|---|---|---|---|---|---|
| Node | Python | PHP | Chrome | Firefox | ||
| Size on disk | 189.4 MB | 396.2 MB | ||||
| mask MAE β | 0.000000 Β· 0.000000 Β· 0.000000 | 0.000001 Β· 0.000009 Β· 0.000024 | ||||
| IoU@0.5 β | 1.00000 Β· 1.00000 Β· 1.00000 | 1.00000 Β· 0.99998 Β· 0.99990 | ||||
| boundary-F (Β±2 px) β | 1.0000 Β· 1.0000 Β· 1.0000 | 1.0000 Β· 1.0000 Β· 1.0000 | ||||
| binary flips β | 0.000 % Β· 0.000 % Β· 0.000 % | 0.000 % Β· 0.001 % Β· 0.003 % | ||||
| Render time CPU, 8k | ~10.1 | ~8.0 | ~8.1 | ~10.6 | ~27.0 | ~26.9 |
| Render time CPU, 4k | ~9.8 | ~7.8 | ~7.9 | ~9.1 | ~26.5 | ~24.5 |
| Render time CPU, 1080p | ~9.8 | ~7.8 | ~8.1 | ~8.4 | ~26.6 | ~23.7 |
| Render time CPU, 512Β² | ~9.8 | ~8.4 | ~8.1 | ~8.3 | ~26.7 | ~23.5 |
| Render time DirectML, 8k | - | ~0.45 | ~0.55 | - | - | - |
| Render time DirectML, 4k | - | ~0.28 | ~0.29 | - | - | - |
| Render time DirectML, 1080p | - | ~0.23 | ~0.22 | - | - | - |
| Render time DirectML, 512Β² | - | ~0.22 | ~0.20 | - | - | - |
| Render time WebGPU, 8k | - | ~0.98 | - | - | ~1.4 | ~4.8 |
| Render time WebGPU, 4k | - | ~0.78 | - | - | ~1.1 | ~3.0 |
| Render time WebGPU, 1080p | - | ~0.68 | - | - | ~0.90 | ~2.5 |
| Render time WebGPU, 512Β² | - | ~0.64 | - | - | ~0.87 | ~2.4 |
| Render time CUDA, 8k | ~0.40 | - | ~0.70 | ~2.5 | - | - |
| Render time CUDA, 4k | ~0.28 | - | ~0.42 | ~1.2 | - | - |
| Render time CUDA, 1080p | ~0.23 | - | ~0.35 | ~0.82 | - | - |
| Render time CUDA, 512Β² | ~0.22 | - | ~0.33 | ~0.65 | - | - |
| Peak memory usage CPU, 8k | ~4.0 GB | ~1.4 GB | ~1.7 GB | ~1.7 GB | ~2.5 GB | ~2.8 GB |
| Peak memory usage CPU, 4k | ~3.8 GB | ~1.4 GB | ~1.5 GB | ~1.5 GB | ~2.3 GB | ~2.7 GB |
| Peak memory usage CPU, 1080p | ~3.8 GB | ~1.4 GB | ~1.4 GB | ~1.4 GB | ~2.3 GB | ~2.4 GB |
| Peak memory usage CPU, 512Β² | ~3.7 GB | ~1.4 GB | ~1.4 GB | ~1.4 GB | ~2.3 GB | ~2.6 GB |
| Peak memory usage DirectML, 8k | - | ~2.2 GB | ~2.2 GB | - | - | - |
| Peak memory usage DirectML, 4k | - | ~2.2 GB | ~2.2 GB | - | - | - |
| Peak memory usage DirectML, 1080p | - | ~2.2 GB | ~2.2 GB | - | - | - |
| Peak memory usage DirectML, 512Β² | - | ~2.2 GB | ~2.2 GB | - | - | - |
| Peak memory usage WebGPU, 8k | - | ~2.0 GB | - | - | ~2.2 GB | ~2.6 GB |
| Peak memory usage WebGPU, 4k | - | ~2.0 GB | - | - | ~2.2 GB | ~2.6 GB |
| Peak memory usage WebGPU, 1080p | - | ~2.0 GB | - | - | ~2.2 GB | ~2.6 GB |
| Peak memory usage WebGPU, 512Β² | - | ~2.0 GB | - | - | ~2.2 GB | ~2.6 GB |
| Peak memory usage CUDA, 8k | ~4.6 GB | - | ~2.4 GB | ~2.4 GB | - | - |
| Peak memory usage CUDA, 4k | ~4.6 GB | - | ~2.4 GB | ~2.4 GB | - | - |
| Peak memory usage CUDA, 1080p | ~4.6 GB | - | ~2.4 GB | ~2.4 GB | - | - |
| Peak memory usage CUDA, 512Β² | ~4.6 GB | - | ~2.4 GB | ~2.4 GB | - | - |
Render time = pre-processing + inference + post-processing in seconds, the median of 5 runs after a warm-up run (session creation excluded). The network always runs at 1024Β², so the input size only changes the host-side resize. Quality rows compare the masks with the original PyTorch model on a fixed 11-image set (product-style samples plus hair / jewelry / portrait stress images), as best Β· mean Β· worst image. Peak memory is the process's peak RAM on CPU (fresh process, session creation + one image; in the browsers, the RAM the page adds); on DirectML / WebGPU / CUDA it is the GPU memory the session adds (Node and every CUDA cell, Safetensors included: that process's own GPU memory; Python's DirectML and the browsers: the whole GPU over the idle level - for the browsers the highest of the four sizes, since the Windows counter under-reads single runs). Python uses the snippet above, PHP onnxruntime-php with GD, the browsers onnxruntime-web with default session options (wasm with 4 threads in a cross-origin-isolated page).
On Chrome the model also runs on WebNN, but slowly: an 8k image takes 10.2 s at 9.7 GB of RAM, against 1.4 s at 2.2 GB on WebGPU.
Firefox is limtied by Bug 1870699 and Bug 1972521 (comment 9), both caused me pain when testing performance of the build.
On CUDA the low-memory graph costs speed. ONNX Runtime's CUDA provider runs it node by node, so the stripes, chunks and DirectML partition cuts of the graph rewrites add up: one 1024Β² inference takes 0.27 s at 2.4 GB of GPU memory, where the earlier plain browser build (389.5 MB) took 0.18 s at 6.5 GB and the PyTorch original takes 0.19 s at 4.6 GB. DirectML compiles the graph and loses no speed to the same rewrites. PHP's CUDA times are mostly GD resizing on the CPU (2.2 of the 2.5 s at 8k).
Tested on
- Hardware - i7-8700K, NVIDIA RTX 5070 Ti 16 GB (driver 591.86), 32 GB DDR4 2400 MT, Windows 11
- Node - Node.js 24.18.0, onnxruntime-node 1.26.0, sharp 0.33.5
- Python - onnxruntime 1.26.0 (CPU), onnxruntime-directml 1.24.4 (DirectML), onnxruntime-gpu 1.26.0 (CUDA, with the pip CUDA 12.9 runtime and cuDNN 9.10.2), numpy 2.3.4, Pillow 12.2.0, Python 3.12.10
- PHP - onnxruntime-php 0.3.7, bundled GD, OPcache JIT off; for CUDA, the GPU
onnxruntime.dllfrom the onnxruntime-gpu 1.26.0 wheel - Browsers - onnxruntime-web 1.26.0, Google Chrome Dev 157.0.8081.0, Firefox Developer Edition 158.0
- fp32 Safetensors - PyTorch 2.12.0 (+cpu, and +cu130 with CUDA 13.0 / cuDNN 9.20 for the CUDA rows), refiners 0.4.1.dev235 (commit 505dbdc), Python 3.12.10
Files
| File | Precision | Size | Notes |
|---|---|---|---|
onnx/model.onnx |
fp32 | ~396 MB | full precision; CPU, WebGPU, DirectML and in the browser (WebGPU / wasm); same masks as PyTorch |
ONNX operator set version 17, ~94.6 M parameters, SHA256(onnx/model.onnx) = a7db4845ed4a557ccb285048e3562c611b762b836d2bd5b809100c34f79a3c1e
Why 396 MB for a 94.6 M-param model: ~378 MB is fp32 weights, the rest is the Swin attention masks and relative-position tables plus the graph itself. The plain export was 805 MB, because constant folding (
do_constant_folding=True) baked ~425 MB of shifted-window masks into it, 405 MB of them all zeros - see Graph rewrites. Disabling the folding makes the plain export larger, not smaller.
I/O contract
| name | shape | dtype | |
|---|---|---|---|
| input | input |
[1, 3, 1024, 1024] |
float32 |
| output | logits |
[1, 1, 1024, 1024] |
float32 (raw logits) |
Pre-processing (must be reproduced by the caller): take the RGB image, resize to 1024Γ1024 (bilinear or a comparable resampler β mild resampler differences don't visibly change the mask), scale to [0,1] (/255), normalize with ImageNet statistics mean = [0.485, 0.456, 0.406], std = [0.229, 0.224, 0.225], layout NCHW.
Post-processing: apply sigmoid to the logits to get an alpha matte in [0,1], then resize it back to the original image size. Use it directly as a mask, or as the alpha channel of an RGBA cutout. (Note: this model emits raw logits β apply sigmoid; do not min-max normalize the output the way some other ONNX matting models, e.g. BEN2, require.)
How it was converted
Exported from the published v0.1 SafeTensors weights with PyTorch's legacy TorchScript exporter (torch.onnx.export with the dynamo path disabled) at opset 17, with constant folding enabled. The raw nn.Module β obtained from refiners' BoxSegmenter and put in eval / float mode β is traced on a single 1Γ3Γ1024Γ1024 float input, and its tensors are named input and logits. There are no architecture changes and no retraining: the exported graph is exactly image β logits.
Two non-obvious choices were required:
- Legacy exporter, not dynamo. torch's dynamo /
torch.exportpath fails decomposing Swin'stranspose(1,2).reshape(...)(it lowers the non-contiguous view to a strictaten.viewthat can't represent the transposed strides). The legacy TorchScript tracer records a realReshapeand exports cleanly. adaptive_avg_pool2dβavg_pool2d. Refiners' pooling always divides evenly (adaptive_avg_pool2d(x, (h//r, w//r))), so it is numerically identical to a plainavg_pool2d(kernel=r, stride=r)that maps to ONNXAveragePoolβ an exact swap, not an approximation.
The graph uses only standard ops (Conv, AveragePool, MatMul, Gemm, Resize, Softmax, LayerNormalization, PRelu, Add, Mul, Concat, Split, Slice, Pad, DepthToSpace, Reshape, Transpose) β no grid_sample / einsum / unfold β so it runs on ORT CPU, WebGPU and DirectML as-is.
Parity (verify_parity.py, CPU fp32): random input β torch vs ORT logits max|Ξ| = 1.5e-5 (sigmoid 3e-6); refiners' golden cactus image β mask MAE = 0.000 / 255 (bit-identical). The full sourced write-up lives in this repo at docs/HOW-CONVERSION-WAS-MADE.md; to reproduce the export yourself, see docs/DEVELOPMENT.md and the python/ scripts.
Graph rewrites (same outputs)
The shipped onnx/model.onnx is the export above after a series of loss-free rewrites. It computes the same function with less memory: against the plain export its logits differ by at most 4.6e-5 on the 11-image eval set (float summation order), with 0 flipped mask pixels; against PyTorch, random-input logits max|Ξ| = 1.7e-5 and the cactus mask is still MAE 0.000 / 255.
- Zero-mask strip - constant folding at export baked ~425 MB of Swin shifted-window masks into the graph, 405 MB of them all-zero pad canvases that each feed a single
ScatterND. They are replaced byConstantOfShape(0), which ONNX Runtime fills at load time. - {-100, 0} masks β fp16 - the 12 remaining additive attention masks hold only -100 and 0, both exact in fp16, so they are stored as fp16 with a
Castback to fp32 (-10 MB). - Static padding - the 24 dynamic-pad
ScatterNDsites, each checked to be bit-identical tonp.pad, become staticPadnodes. Without this, onnxruntime-web's constant folding of their int64 index arithmetic needs ~2.6 GB and session creation aborts in the 4 GB wasm heap. - Mask broadcast - the window-mask
Expandis dropped andAddbroadcasts the mask instead (folding thatExpandadded another ~0.9 GB at load in the browser). - Low-memory rewrite - ~95% of the runtime memory is activations, and ONNX Runtime keeps a dead buffer until the next tensor of the same shape takes it over. The full-resolution
shallowconv is folded into strided convs on the input image, every nearest-upsample + 3Γ3 conv becomes four 2Γ2 "polyphase" convs +DepthToSpace, the 512Β²/1024Β² tail runs in 16 row stripes, the Swin backbone runs once per view (5 Γ batch 1, shared weights) stage by stage, both 128Β² cross-attention blocks run in 8 token chunks, a 128β128 no-opResizeis dropped and the static shape arithmetic is folded into constants. - DirectML partition cuts - DirectML compiles connected runs of nodes into its own graphs and plans their memory itself, running independent branches side by side without reusing their buffers, so step 5 alone does nothing there. 166 no-op
Concat(x, Slice(x, 0:0))pairs make every view, chunk and stripe a separate DirectML graph. - WebGPU fan-out - WebGPU binds every input and output of a
Concat/Splitas one storage buffer, and the default limit is 8 per shader. Steps 5-6 added wider ops (16 stripes, 8 chunks, 5 views); they are split into two-level trees of at most 4, so no op needs more than 5 buffers, as in the plain export.
Steps 1-4 take the file from 805 to 389.5 MB and make it load in browsers; steps 5-7 add 6.8 MB and cut peak memory (this machine, see Comparative numbers) on CPU from 2.3 to 1.4 GB, on WebGPU from 5.8 to 2.0 GB and on DirectML from 7.7 to 2.2 GB, at about the same speed, with only the session creation taking a few seconds longer.
Quality
Metrics were scored with PySODMetrics in two modes, which measure different things.
| Mode (same 120 masks) | MAE β | S-measure β | E-measure (mean) β | Dice (mean) β |
|---|---|---|---|---|
| Box-prompted crop β the base model's published protocol | 0.0079 | 0.9738 | 0.9854 | 0.9669 |
| Whole-image (no box) β how this release is used | 0.086 | 0.796 | 0.793 | 0.712 |
Box-prompted crop mode confirms the conversion: bit-identical parity implies identical task quality, and running the base model's own protocol on this ONNX reproduces its published numbers to within rounding:
| Metric | ONNX | Official Safetensors | Ξ |
|---|---|---|---|
| MAE β | 0.0079 | 0.0078 | +0.0001 |
| S-measure β | 0.9738 | 0.974 | β0.0002 |
| E-measure (mean) β | 0.9854 | 0.985 | +0.0004 |
| Dice (mean) β | 0.9669 | 0.967 | β0.0001 |
n = 120. The sub-0.001 deltas are host-side resize / box-derivation rounding, not conversion error β the ONNX inherits the original's quality (which, on this set, beats briaai/RMBG-1.4 and box-guided ZhengPeng7/BiRefNet; see the base card).
Whole-image mode (no box) scores far lower by construction: this dataset's ground truth is one specific boxed object, so scoring the no-prompt salient output against it penalizes the model whenever the image's salient region isn't exactly that boxed target. With this data I wanted to show a worst-case lower bound on a box-oriented set , it is not representative of ordinary single-subject background removal (see the Examples above for what whole-image output actually looks like).
Usage β Transformers.js (easiest)
This repo ships the standard onnx/model.onnx + config.json + preprocessor_config.json layout, so Transformers.js drives the whole pipeline for you β pre-process, sigmoid, and resize-back included:
import { pipeline } from '@huggingface/transformers';
const segmenter = await pipeline('background-removal', 'MarcinEU/finegrain-box-segmenter-ONNX');
const out = await segmenter(['https://example.com/photo.jpg']);
out[0].save('cutout.png'); // RawImage RGBA - alpha is the matte
// out[0].toCanvas() / await out[0].toBlob() if you'd rather not save to disk
Verified end-to-end (@huggingface/transformers v4.2.0): the model resolves to SwinForSemanticSegmentation, Transformers.js remaps the processor output onto our input tensor and auto-applies sigmoid to the raw logits β no custom code. Add { device: 'webgpu' } for the GPU. Use v4.2.0 or newer (the automatic input-name remap and auto-sigmoid are required). If you need full control over pre/post-processing, use the hand-rolled paths below.
Usage β onnxruntime-web (browser, WebGPU)
import * as ort from 'onnxruntime-web/webgpu';
const SIZE = 1024, MEAN = [0.485,0.456,0.406], STD = [0.229,0.224,0.225];
const session = await ort.InferenceSession.create(
'https://huggingface.co/MarcinEU/finegrain-box-segmenter-ONNX/resolve/main/onnx/model.onnx',
{ executionProviders: ['webgpu'] });
// draw your image into a 1024x1024 canvas, then:
const { data } = ctx.getImageData(0, 0, SIZE, SIZE); // RGBA, Uint8ClampedArray
const plane = SIZE * SIZE, chw = new Float32Array(3 * plane);
for (let p = 0; p < plane; p++) {
chw[p] = (data[p*4] / 255 - MEAN[0]) / STD[0];
chw[plane + p] = (data[p*4+1] / 255 - MEAN[1]) / STD[1];
chw[2*plane + p] = (data[p*4+2] / 255 - MEAN[2]) / STD[2];
}
const out = await session.run({ input: new ort.Tensor('float32', chw, [1,3,SIZE,SIZE]) });
const logits = out.logits.data; // Float32Array, 1024*1024
const alpha = logits.map(v => 1 / (1 + Math.exp(-v))); // sigmoid -> [0,1]
// resize `alpha` back to your image size and composite as the alpha channel.
Usage β onnxruntime-node (Node.js, with sharp)
A complete, runnable background remover is in usage/remove_bg.mjs β run it from the repo root (with onnx/model.onnx downloaded):
npm i onnxruntime-node sharp
node usage/remove_bg.mjs --image photo.jpg --model onnx/model.onnx # CPU
node usage/remove_bg.mjs --image photo.jpg --model onnx/model.onnx --ep webgpu # GPU (~10x faster here; see Comparative numbers below)
It resizes the whole image to 1024Β², normalizes, runs the session, applies sigmoid, resizes the mask back, and writes *_mask.png and *_cutout.png (RGBA). Outputs carry real transparency β view them over a checkerboard/solid background.
Usage β Python (onnxruntime)
Verified against this exact export (pip install onnxruntime pillow numpy huggingface_hub); masks match the Node reference pipeline to ~1/255 mean difference:
from huggingface_hub import hf_hub_download
from PIL import Image
import numpy as np, onnxruntime as ort
model_path = hf_hub_download("MarcinEU/finegrain-box-segmenter-ONNX", "onnx/model.onnx")
session = ort.InferenceSession(model_path)
img = Image.open("photo.jpg").convert("RGB")
x = np.asarray(img.resize((1024, 1024), Image.BILINEAR), dtype=np.float32) / 255.0
x = ((x - [0.485, 0.456, 0.406]) / [0.229, 0.224, 0.225]).transpose(2, 0, 1)[None].astype(np.float32)
logits = session.run(None, {"input": x})[0][0, 0]
alpha = 1.0 / (1.0 + np.exp(-logits)) # sigmoid -> [0,1]
mask = Image.fromarray((alpha * 255).astype(np.uint8)).resize(img.size, Image.BILINEAR)
cutout = img.copy(); cutout.putalpha(mask)
cutout.save("cutout.png"); mask.save("mask.png")
For NVIDIA CUDA, pip install onnxruntime-gpu[cuda,cudnn], call ort.preload_dlls() before creating the session (it loads the CUDA and cuDNN DLLs those extras install) and pass providers=["CUDAExecutionProvider", "CPUExecutionProvider"]. One catch: the newest cuDNN on pip (9.27) made every convolution of ORT 1.26 fail here with CUDNN_BACKEND_API_FAILED, and ORT then quietly re-runs the model on the CPU; 9.10.2 works (pip install nvidia-cudnn-cu12==9.10.2.21).
Notes
- Resolution / batch are fixed at
1Γ3Γ1024Γ1024(the network hard-codes 1024Β² internally). Any aspect ratio works β the squash-resize matches the original model's behavior (a very common approach among mask-generation and image-segmentation models). - Runtime requirement: the graph is opset 17, so it needs
onnxruntime-nodeβ₯ 1.12.0 (July 2022; earlier versions will crash on load). Currentonnxruntime-webis fine. - WebGPU: runs as-is on standard adapters β the widest
Concat/Splithas 4 inputs/outputs (5 storage buffers), under the defaultmaxStorageBuffersPerShaderStagelimit of 8 (step 7 of the graph rewrites keeps it that way). Only WebGPU "compatibility mode" (limit 4) devices would need a further cascade rewrite.
License & attribution
MIT, inherited from the base model. All credit for the weights and architecture goes to Finegrain (finegrain/finegrain-box-segmenter) and the MVANet authors. This repository only provides an ONNX conversion of the public v0.1 weights (model.safetensors, SHA256 fd5f13919dfc0dda102df1af648c3773c61221aa65fe58d6af978637baded1ae).
Citation
@article{mvanet,
title = {Multi-view Aggregation Network for Dichotomous Image Segmentation},
author = {Yu, Qian and Zhao, Xiaoqi and Pang, Youwei and Zhang, Lihe and Lu, Huchuan},
journal = {CVPR},
year = {2024}
}
- Downloads last month
- 18
Model tree for MarcinEU/finegrain-box-segmenter-ONNX
Base model
finegrain/finegrain-box-segmenter















































































// npm i @huggingface/transformers import { pipeline } from '@huggingface/transformers'; // Allocate pipeline const pipe = await pipeline('mask-generation', 'MarcinEU/finegrain-box-segmenter-ONNX');