• Coriza@lemmy.world
    link
    fedilink
    English
    arrow-up
    2
    ·
    20 hours ago

    I am not sure it would not help commercial solutions, if all experts are used all the time, sure, but if for example the usage is biased for some experts it would enable one machine to serve more users in parallel or save on VRAM or DRAM without compromising response time, hence cutting costs.

    • KingRandomGuy@lemmy.world
      link
      fedilink
      English
      arrow-up
      1
      ·
      4 hours ago

      It still won’t help for a couple of reasons. For one, VRAM is fairly abundant on commercial deployments. Even for big models, a company is probably deploying on 1-2 nodes of 8x H200 or newer (hence several TB of VRAM).

      But more importantly, commercial inference relies on heavy concurrency. So even if some experts are uncommon, with a lot of concurrent users, they will still fire frequently enough for the performance difference to be felt. And in my own experience, expert use isn’t uniform but it isn’t particularly biased either. This is especially tough since high-concurrency inference can actually be fairly compute bound, but this expert caching system either starves the system of bandwidth (if you require compute on the GPU, then you’re stuck with PCIe speeds which are tiny compared to HBM and even DRAM), or you’re starved of compute (if you do compute on the CPU).

      It’s nice for local inference, but yeah, not representative of commercial inference.