<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Yolo Ultralytics stopped exporting fp16 tflite format. Can tflite-server directly handle fp32.tflite models?]]></title><description><![CDATA[<p dir="auto">Hey ModalAI,</p>
<p dir="auto">Your documentation on custom inference using voxl-tflite-server is quite outdated. Yolo Ultralytics has stopped exporting frozen graph model format to tflite fp16 quantization. Ultralytics says there's no quantize=16 option — an FP32 .tflite model automatically runs in FP16 at runtime when using a GPU delegate (WebGPU/OpenCL/Metal), so a separate FP16 export file isn't needed. That's why half=True/quantize=16 throws the assertion.</p>
<p dir="auto">Can voxl-tflite-server automatically execute FP32 models internally in FP16?<br />
Hardware: VOXL2<br />
Source code: Latest voxl-tflite-server code on modalai gitlab page.</p>
]]></description><link>https://forum.modalai.com/topic/5356/yolo-ultralytics-stopped-exporting-fp16-tflite-format.-can-tflite-server-directly-handle-fp32.tflite-models</link><generator>RSS for Node</generator><lastBuildDate>Thu, 30 Jul 2026 21:38:41 GMT</lastBuildDate><atom:link href="https://forum.modalai.com/topic/5356.rss" rel="self" type="application/rss+xml"/><pubDate>Thu, 30 Jul 2026 16:44:48 GMT</pubDate><ttl>60</ttl></channel></rss>