<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Alibaba's LLM Qwen2-0.5B for VOXL 2]]></title><description><![CDATA[<p dir="auto">hi,<br />
Has anyone tried Qwen2-0.5B  for the VOXL 2?</p>
<p dir="auto">Here's are some requirements and compatibility:</p>
<ol>
<li>RAM Requirements for Qwen2-0.5B</li>
</ol>
<blockquote>
<blockquote>
<p dir="auto">Base model (FP16): ~1GB of RAM<br />
Quantized to INT8: ~500MB of RAM<br />
Quantized to INT4: ~250MB of RAM</p>
</blockquote>
</blockquote>
<p dir="auto">VOXL 2 Specifications</p>
<blockquote>
<blockquote>
<p dir="auto">The VOXL 2 has approximately 4GB of RAM available<br />
It uses a Qualcomm QRB5165 processor<br />
Has AI acceleration capabilities via the Hexagon DSP<br />
Primarily designed for computer vision tasks</p>
</blockquote>
</blockquote>
<p dir="auto">Compatibility Assessment<br />
Qwen2-0.5B could potentially run on the VOXL 2 with quantization to INT4 or INT8, the memory footprint would be manageable.</p>
<blockquote>
<blockquote>
<p dir="auto">Inference speed would likely be slow, perhaps 1-2 seconds per token</p>
</blockquote>
</blockquote>
<p dir="auto">Question:</p>
<p dir="auto">Has ModalAI explored using TensorFlow Lite with the Hexagon delegate to run compact language models like Qwen2-0.5B on the VOXL 2 platform? I'm interested in whether you've experimented with leveraging the Qualcomm GPU for language model inference, even though I understand the VOXL 2 is primarily optimized for computer vision workloads. Have you conducted any experiments or performance testing with small LLMs on this hardware?<br />
Thanks.<br />
suvasis</p>
]]></description><link>https://forum.modalai.com/topic/4373/alibaba-s-llm-qwen2-0-5b-for-voxl-2</link><generator>RSS for Node</generator><lastBuildDate>Tue, 08 Sep 2026 16:09:34 GMT</lastBuildDate><atom:link href="https://forum.modalai.com/topic/4373.rss" rel="self" type="application/rss+xml"/><pubDate>Thu, 24 Apr 2025 04:17:37 GMT</pubDate><ttl>60</ttl></channel></rss>