Skip to content

Even more artificial intelligence

Barber Shop Talk
10 3 545
  • I'm curious Facebook has advertisements out the ass for artificial intelligent chatbots that can do everything from predict the weather to giving car repair instruction to researching 15th century British poetry to looking up cast members in movies too handicapping a sports event to diagnosing medical and psychological problems and giving virtual blow job's.

    Is anybody here familiar with this phenomenon, has anybody tried one is anybody going to invite the bronze board to your wedding with one?

    WSS 

  • Not sure if this has been in the MSM, but this is really big deal.

    image.thumb.jpeg.93a2e525c34f988c484d0bc86ffcfb81.jpeg

    Deepseek-r1 was just released to the public late last week.

    It's blown up since then, I've seen all kinds of positive posts about it.

    Got me curious, as this one you can set up to run locally.

    I set up a VM (lowers performance) on an external backup drive (even more performance degradation).

    I did it this way just to have more isolation that won't impact my Computer while I'm doing other things while it runs in the background.

    There are currently 7 different data models available from 1.1 GB up to 404 GB, most are "distilled" from the original, so there will be some accuracy issues due to that.

    I'm fooling around with the 4.7 GB and 9.0 GB models, and while being slowed due how I set things up, it's pretty interesting.

    I can easily see how this could be an assistive program for many things, from research to programming.

  • image.png.fb99e069d08721905e2f26f66d9c02a0.png

    image.thumb.jpeg.b7f4f623170871643176ac780578335f.jpeg

    image.thumb.jpeg.971a1d382b3e89d678bbe256388a4d24.jpeg

    I can attest to seeing this behavior in several of the queries I've made.

    I've also seen it suggest additional info to add to queries as it evaluates so that it can give better / more accurate answers.

    For some queries on programming, I've had it add several suggestions for additional features and some modules to link and how to set them up for more functionality.

    Not saying it's perfect, but that may just be the way the query was structured.

    Pretty impressive though so far.

    Could be scary too (Matrix / Terminator).

    There have been about five new AI models from different companies come out in the last week.

    I saw one report where someone had the AI rewrite a plugin it used and it doubled the performance for it.

    I've not seen anything like that, though I'm not that deeply into it yet, nor sure if I will be.

  • Some of the hype about how fast this seems to be ramping up.

    image.png.be32db13e0106e89c23d9515ba286236.png

  • I heard an expert say, that it was copied nearly 95%.....from open source chat gpt.

    AI concerns me - like the belligerent misuse of any tool is a concern.

    AI can end up becoming a "black hole" that sucks up information in certain areas, like

    military information, defense contractor personnel files......Technology blueprints.....

    for a while, I worked in a mail room in the service - loved it. One day, a B-52 pilot came in, carrying a blue folder, got his mail, had to leave for a meeting. I turned around a minute later, and his blue folder was laying there - he had put all his mail on top of it, and it slipped from being picked up and he didn't notice.

       I looked at it, to get his name, and saw it was "classified". It contained wheel bearing schedules for the B-52 - based on the number of landings I think it was.

       Little things like that - could give an enemy an advantage if they knew it. At a remote location, if they knew in a war that the planes were needing to be serviced at a certain point, they could hit that base while the planes were in maintence for a day.

       So, "loose lips sink ships", and AI could be gleaning important small details that could be used to an enemies advantage.

       AI in the medical field- awesome. But on the flip side, any tool can be used for nefarious purposes. That is scary.

  • There has been about one new AI model released a week (different Companies), this is the latest (Grok 3 - came out late Monday) :

    image.thumb.png.e4ee26281fd012960dfcc9780879c6af.png

    Grok 3 is Elon's project   .   .   .   there are hints that more is going on with this project than just improved processing capacity / speed.

    One too watch.

  • This one came out last week.

    image.png.93673a967f7a87189e3b2479b200ccef.png

    Non trivial stuff.

  • A bit over a year later and just an (amazing) update on where things have progressed too since Jan/Feb 2025.

    Based on updated metrics :

    image.png.44150726b0eb85e8b79e18c7444ba3b8.png

    Recall on previous posts (above) this thread how ground breaking this model was, and I'm talking about the 8B (8 Billion) parameter model, this rating is on the 528B (685B ?) parameter model.

    19.

    These are the current (as of a few days ago) rankings (it's outdated from things I've seen) : 

    image.png.df8e3ab89390bd0a4e1779107046f46b.png

    Note, the Deepseek on the above graph is Deepseek V3.2, not the R1 from last Jan/Feb.

    The above are the best of the subscription models ($).

    I've been running this one locally (27B) as that is the best I can do based on INT and VM memory available :

    (These are all free at 40B or less for local running by the way).

    image.png.8389f381ca26a220c2eb2b53a71ca2eb.png

    27B is very slow (VM only, no GPU processing) on my system, but has been extremely impressive.

    Still, at more than twice the INT of the Deepseek R1 528B model it is quite tolerable for what I'm testing on it.

    The 9B model is quite a bit quicker, and while the "Intelligence" rating is a bit lower than the 27B, it is still quite amazing, and far outperforms the Deepseek 528B model from a little over a year ago.

    Just for reference since I posted the subscription models, these are the current best of the free (local if you can run them) models :

    image.thumb.png.e92db7be6cf7be4c306df8e8781f4f6c.png

    Note, the above are sorted more by parameters than intelligence rating - just the way the site is set up.

    So 57 INT for the best subscription model vs 54 INT for the best Open source model - it's getting incredibly close providing you have the resources.

    I should also add the caveat that some models perform better at certain tasks than others, and that the above is a composite overall assessment.

    I should note that there are also hybrids, meaning a mix of two or more models into one model.

    The latest I've seen is a mix of QWEN3.5 (9B) & GLM5.1 (9B) (distilled models) coming in at 18B, which supposedly outperforms QWEN3.5 27B (3.5 27B is better than 3.5 35B surprisingly).

    QWEN 3.6 35B was just released yesterday or the day before and is on par or marginally better than QWEN 3.5 27B.

    Expecting QWEN 3.6 27B to be markedly better, but may be a week or two out yet.

    Current best moderate requirement free model looks to be the hybrid QWEN/GLM 18B, which is actually mind boggling.

    Once QWEN 3.6 27B comes out, I expect that to be the king for free moderate capacity systems.

    QWEN 3.6 9B should be a week or two behind 27B, but I expect it to be extremely performant, possibly twice the level of Deepseek R1 528B from just over a year ago. 

    There has been some incredible progress made in the last year on these systems.

    I believe Anthropic, and maybe others have turned over (self) programming their models to the AI itself (scary) - I may have more on this later.

    If this is truly this case, i would expect the pace of progress on these models to increase beyond what they have already over the past year.

  • image.thumb.png.f034ad3daeab01e722f2c1e2deaf98b5.png

    Anyone getting concerned yet ?

  • Surprised QWEN 3.6 27B came out yesterday, and as expected it was a decent improvement :

    image.png.e6c92509f5ace2ecf99d7e3e8e96835f.png

    From INT 42 released a couple months ago to INT 46, so nearly a 10% bump.

    Should see QWEN 3.6 9B next week based on the rate of these recent releases.