There has been an overwhelming amount of new models hitting HuggingFace. I wanted to kick off a thread and see what open-source LLM has been your new daily driver?

Personally, I am using many Mistral/Mixtral models and a few random OpenHermes fine-tunes for flavor. I was also pleasantly surprised by some of the DeepSeek models. Those were fun to test.

I believe 2024 is the year open-source LLMs will catchup with GPT-3.5 and GPT-4. We’re already most of the way there. Curious to hear what new contenders are on the block and how others feel about their performance/precision compared to other state-of-the-art (closed) source models.

    • Blaed@lemmy.worldOPM
      link
      fedilink
      English
      arrow-up
      2
      ·
      2 years ago

      What sort of tokens per second are you seeing with your hardware? Mind sharing some notes on what you’re running there? Super curious!

  • Frozen_byte@sffa.community
    link
    fedilink
    English
    arrow-up
    7
    ·
    2 years ago

    I would also be interested in Code-Pilot Models that are reaching for same performance like GitHub or Microsofts paid Models.

    Currently I use TabbyML but the available Models are by far inferior.

      • Blaed@lemmy.worldOPM
        link
        fedilink
        English
        arrow-up
        3
        ·
        2 years ago

        I was pleasantly surprised by many models of the Deepseek family. Verbose, but in a good way? At least that was my experience. Love to see it mentioned here.

  • 🇨🇦Samuel Proulx🇨🇦@rblind.com
    link
    fedilink
    English
    arrow-up
    2
    ·
    2 years ago

    Personally I find myself renting GPU and running Goliath 120b. Smaller models could do what I’m doing if I spent more time optimizing my prompts. But every day I’m doing different tasks, and Goliath 120b will just handle whatever I throw at it, no matter how sloppy I am. I’ve also been playing with LLAVA and Hermes vision models to describe images to me. However, when I really need alt-text for an image I can’t see, I still find myself resorting to GPT4; the open source options just aren’t as accurate or detailed.

  • wilkinsonwilfrid
    link
    fedilink
    English
    arrow-up
    1
    ·
    edit-2
    1 year ago

    @slope game Several Deepseek models pleasantly surprised me. Flowing, yet with an air of elegance? I can say that from personal experience. Delighted to see it brought up here.