HN user

mlashuel

12 karma
Posts10
Comments7
View on HN

Hey I built an AI Copilot on ChatGPT that understands your page content.

This helps with two things.

1. You don't have to waste time going back and forth between chatGPT and your main tab.

2. Because our copilot understands your content you don't have to copy and paste information from your browser to chatgpt, to ask AI about what you are observing. Our copilot is context aware and can help whenever you need it too.

Looking forward for feedback. Browsercopilot is free to use for now.

[dead] 2 years ago

Imagine taking the unique selling points of each LLM and combining them into one.

Ex. GPT-4o is more logical, where as Claude is more creative, this combines both unique selling points of each model.

When we spoke to our customers they asked if we could synthesize all the outputs from all the LLMs into one high quality response.

So Today, we are launching Mixture AI.

Mixture AI (MoA) is a novel approach that leverages the collective strengths of multiple LLMs to enhance performance, achieving state-of-the-art results.

By employing a layered architecture where each layer comprises several LLM agents, MoA significantly outperforms GPT-4 Omni’s 57.5% on AlpacaEval 2.0 with a score of 77.1%!

Give it a shot on ChatPlayground AI

[dead] 2 years ago

GPT4o - 93.85% Claude Opus - 92.31% Gemini 1.5 Pro - 90.77% Meta.ai - 90.77% Wizard LM2 8x22b - 90.77% Claude Sonnet - 90.77% GPT4 - 89.23% Qwen 1.5 100B - 89.23% Phind - 87.69% Qwen Max - 87.69% Mixtral 8X22B - 87.69%

Hi everyone,

I recently conducted a test to compare the performance of 12 different AI models. Using a custom-made 65-question exam, I evaluated each model's accuracy and compiled the results. I thought it would be interesting to share the findings here and get your thoughts on them.

I was really surprised by some of these results. All the models did quite well, considering the questions were pretty challenging. Even the lowest scores were still at 87.69%. It was fascinating to see open-source models being very competitive with the best models available.

I bought the lifetime deal on chatarena.ai to have an interface to quickly test all of them.

Also, I tested Perplexity Pro, but it performed so poorly that I didn't include it. Maybe the 65-question test was too long for it? Usually, it gives me pretty good results, so I'm not sure what happened there.

In terms of general personal use outside this test, currently, GPT-4, GPT-4o, and Gemini 1.5 are my favorites. I'm still not sure which is better between GPT-4 and GPT-4o as the outputs tend to be quite similar. I used to be really disappointed with Gemini, but I think they've improved it a lot recently. On the other hand, I'm more disappointed with Claude, especially Opus. Yes, it did rank 2nd here, but it constantly messes up in my personal use where GPT and Gemini never do. I let my friends use my subscriptions sometimes, and they agree that GPT and Gemini have been outperforming Opus.

Please share your insights and thoughts on these results!

[dead] 2 years ago

I compared 10 AI Chatbots, the chatbots i compared and the questions I asked are all at the bottom of this post.

It is very important to note this research was done on a very small dataset of 22 questions.

The best AI Chatbot answer for each question was decided by me simply putting myself in the choose of the person asking that question and deciding which output is the most helpful.

Here are my main takeaways.

No one should still be using GPT 3.5 or Gemini 1.0, They may get the job done, but 98% of the time one of these AI Chatbots (Gemini 1.5 Pro, Bing Copilot, ChatGPT4o, and Perplexity) will give you a better response.

The best way to always get the best response is by always comparing multiple AI Chatbots and choosing the best answer. This is because

If you are only using ChatGPT-3.5, llama3, claude sonnet, or mistral 100% of the time there is another AI Chatbot with a better answer.

If you are only using Gemini 1.0 96% of the time there is another AI Chatbot with a better answer.

If you are only using perplexity 86% of the time there is another AI Chatbot with a better answer.

If you are only using ChatGPT-4o 73% of the time there is another AI Chatbot with a better answer.

If you are only using Gemini 1.5 Pro 73% of the time there is another AI Chatbot with a better answer.

If you are only using Bing Copilot 73% of the time there is another AI Chatbot with a better answer.

I personally used chatplayground.ai to compare all of them.

I stopped using chatgpt 3.5 a long time ago because this gives you all the Premium AI Chatbots for the same price as GPT-4.

___Notes on each Chatbot____

-- Perplexity is good at giving you a completely answer, most AI Chatbots will give you bullet point answers, perplexity perfers to write out a complete answer rather than just giving bullet points

-- Bing Copilot is great because it cites its resources, will even recommend videos to help you

-- ChatGPT-4o has really descriptive and usually longer answers, but prefers to answer in bullet points

-- Gemini 1.5 pro is great because it feels like it understands the context of your question more by having a conversational tone

-- Bing copilot can be great but 20-30% of the time the answers it gives are not even usable

-- Because Bing copilot is using sources to give you answers the answers feel much more human like and at times more useful then other AI Chatbots that just list a bunch of basic bullet points

-- Their should be no reason anyone is still using ChatGPT 3.5 today

-- Gemini 1.0 gives good answers but not as detailed and helpful as Gemini 1.5 Pro

-- ChatGPT-4o was able to generate really nice data tables, that Gemini 1.5 Pro wasn't able to.

___ChatBots____

ChatGPT-3.5

ChatGPT-4o

Gemini 1.0

Gemini 1.5 Pro

Bing Copilot

Claude Sonnet

Llama 3

Mixtral 8x7b

Mistral Large

Perplexity

Hey Everyone

As you guys all now loom is a very popular internal feedback tool. I always wanted to use loom collect feedback from my users but there was no way to automate that process.

So I created a simple widget you can install in your website in 5 minutes, that allows your visitors to submit loom video feedback and annotated screenshots without requiring an account.

[dead] 3 years ago

If you go to Sam Altmans Product hunt profile it shows he launch a new product called Redoc. Basically like a chatgpt but for copywriting. Thoughts on this new product I tried it and it’s pretty neat. But I’m confused why it hasn’t been announced, is it supposed to be like a gem.