HN user

Gcam

87 karma

twitter: https://twitter.com/grmcameron

Posts9
Comments29
View on HN

Artificial Analysis | https://artificialanalysis.ai | Full Stack Engineer, ML Engineer, Member of Technical Staff | Onsite in San Francisco or in Australia/New Zealand | Competitive Salary + Equity Artificial Analysis is an independent AI benchmarking and insights provider. We benchmark AI to help engineers and companies understand AI and make informed decisions regarding which AI technologies to use. We are fast growing with a team of 20 and have backing from investors including Nat Friedman, Daniel Gross & Andrew Ng.

We are hiring for four roles:

1. Full Stack Engineer: Full Stack Engineer to support our benchmarking of AI and with communicating these benchmarks to our users. Proficiency in Typescript & Python required. Familiarity with LLM APIs preferred. Tech stack: Javascript/Typescript, Node.js, React/Next.js, Python.

2. ML Engineer: ML Engineer to support our benchmarking and evaluation of AI software stack. You will design and run benchmarks and evaluations of different AI models. Strong analytical skills and proficiency in Python required.

3. Member of Technical Staff: You will develop benchmarks to test AI models and work to translate technical insights to analysis that helps companies navigate AI.

Apply at hiring (-at-) artificialanalysis.ai with your resume, github and dot points on relevant experience (including anything you've built). Add | HackerNews to email subject line.

Artificial Analysis | https://artificialanalysis.ai | Full Stack Engineer, ML Engineer, Member of Technical Staff | Onsite in San Francisco or in Australia/New Zealand | Competitive Salary + Equity

Artificial Analysis is an independent AI benchmarking and insights provider. We benchmark AI to help engineers and companies understand AI and make informed decisions regarding which AI technologies to use. We are fast growing with a team of 20 and have backing from investors including Nat Friedman, Daniel Gross & Andrew Ng.

We are hiring for four roles: 1. Full Stack Engineer: Full Stack Engineer to support our benchmarking of AI and with communicating these benchmarks to our users. Proficiency in Typescript & Python required. Familiarity with LLM APIs preferred. Tech stack: Javascript/Typescript, Node.js, React/Next.js, Python.

2. ML Engineer: ML Engineer to support our benchmarking and evaluation of AI software stack. You will design and run benchmarks and evaluations of different AI models. Strong analytical skills and proficiency in Python required.

3. Member of Technical Staff: You will design evaluations of different AI models and work to translate technical insights to analysis that helps companies navigate AI.

4. Product Manager (AI Media Generation): You'll work closely with us to develop and enhance our media generation (Image, Video, Speech, Music) arenas and leaderboards, contributing to product strategy and execution at the intersection of creativity and AI.

Apply at hiring (-at-) artificialanalysis.ai with your resume, github and dot points on relevant experience (including anything you've built). Add | HackerNews to email subject line.

Artificial Analysis | https://artificialanalysis.ai | Full Stack Engineer & ML Engineer | Onsite in San Francisco or in Australia/New Zealand | Competitive Salary + Equity

Artificial Analysis is an independent AI benchmarking and insights provider. Our benchmarks help engineers and companies understand AI and make informed decisions on AI technologies.

We are hiring for two roles:

1. Full Stack Engineer: Full Stack Engineer to support our benchmarking of AI and with communicating these benchmarks to our users. Proficiency in Typescript & Python required. Familiarity with LLM APIs preferred.

Tech stack: Javascript/Typescript, Node.js, React/Next.js, Python.

2. ML Engineer: ML Engineer to support our benchmarking and evaluation of AI software stack. You will design and run benchmarks and evaluations of different AI models.

Strong analytical skills and proficiency in Python required.

Apply at hiring (-at-) artificialanalysis.ai with your resume, github and dot points on relevant experience.

Artificial Analysis | https://artificialanalysis.ai | Full Stack Engineer & ML Engineer | Onsite in San Francisco or in Australia/New Zealand | Competitive Salary + Equity

Artificial Analysis is an independent AI benchmarking and insights provider. Our benchmarks help engineers and companies understand AI and make informed decisions on AI technologies and providers.

We are hiring for two roles:

1. Full Stack Engineer: Full Stack Engineer to support our benchmarking of AI and with communicating these benchmarks to our users. Proficiency in Typescript & Python required. Familiarity with LLM APIs preferred.

Tech stack: Javascript/Typescript, Node.js, React/Next.js, Python.

2. ML Engineer: ML Engineer to support our benchmarking and evaluation of AI software stack. You will design and run benchmarks and evaluations of different AI models.

Strong analytical skills and proficiency in Python required.

Apply at hiring (-at-) artificialanalysis.ai with your resume, github and dot points on relevant experience.

Artificial Analysis | https://artificialanalysis.ai | Senior Full Stack Software Engineer | Onsite in San Francisco | Competitive salary + Equity

Artificial Analysis is an independent benchmarking, evaluation and insights provider for AI. Our benchmarks let engineers and companies make the best decisions on which technologies and providers to use, empowering them to build the next generation of AI applications.

We are looking for a full stack developer to support us in our analysis of AI and presenting it to the world at https://artificialanalysis.ai/.

Strong analytical skills and proficiency in Typescript & Python required. Familiarity with LLMs and AI scaling laws preferred.

Tech stack: Javascript/Typescript, Node.js, React/Next.js, Python.

Apply here: https://artificialanalysis.ai/careers

Artificial Analysis | https://artificialanalysis.ai | Senior AI Analyst & Senior Software Developers | Remote or Hybrid (USA, San Francisco preferred, or Australia) | Competitive salary + Equity

We're seeking a Senior AI Research Analyst to support with benchmarking and evaluation of AI. Role involves analyzing AI systems, visualizing data, and supporting people in understanding the capabilities of AI (across different modalities).

Strong analytical skills, AI/ML research experience, and proficiency in Python/data analysis required (Typescript a nice to have). Familiarity with LLMs and AI scaling laws preferred.

Apply here: https://artificialanalysis.ai/careers

Model quality index methodology is as per this comment (can add perplexity using the dropdown): https://news.ycombinator.com/item?id=39014985#39017632

It's a combination of different quality metrics which have Perplexity, overall, not performing as well. That being said, I think we are in the very early stages of model quality scoring/ranking - and (for closed sourced models) we are seeing frequent changes. Will be interesting to see how measures evolve / model ranks change

Hi HN, Thanks for checking this out! Goal with this project is to provide objective benchmarks and analysis of LLM AI models and API hosting providers to compare which to use in your next (or current) project. Benchmark comparisons include quality, price, technical performance (e.g. throughput, latency).

Twitter thread with initial insights: https://twitter.com/ArtificialAnlys/status/17472648324397343...

All feedback is welcome

We have this (and other more detailed metrics) on the models page https://artificialanalysis.ai/models if you scroll down and for individual hosts if you click into a model (nav or click one of the model bars/bubbles) :)

There are some interesting views of throughput vs. latency whereby some models are slower to the first chunk but faster for subsequent chunks and vice versa, and so suit different use cases (e.g. if just want a true/false vs. more detailed model responses)

Quality index is equally-weighted normalized values of Chatbot Arena Elo Score, MMLU, and MT Bench.

We have a bit more information in the FAQ: https://artificialanalysis.ai/faq but thanks for the feedback, will look into expanding more on how the normalization works. We are thinking of ways to improve this generalized metric.

A sticking point is quality can of course be thought of from different perspectives, reasoning, knowledge (retrieval), use-case specific (coding, math, readability), etc. This is why show individual scores on home page and models page: https://artificialanalysis.ai/models

Thanks for the feedback! Yes, agree this would be a good idea. We don't have this view but best place to get an idea of this with current site would be the /models page (https://artificialanalysis.ai/models) and scrolling to the over time graphs and looking at the variance. To see if being driven by individual hosts can also click into the by-model pages and see the over time graphs, e.g. https://artificialanalysis.ai/models/mixtral-8x7b-instruct

Thanks for the letting me know. Odd as not occurring with my iOS Safari, can anyone else please let me know if they are encountering this issue (any their iOS version if possible). There is a console error but should be just a react defaultprops deprecation notice from a library being used (should not break DOM)

We [0] took the same approach since launching (which has been a while now). We focused on: 1) Outbound calls/emails, and 2) Networks

Rather than relying on the ROI of long term marketing techniques such as content marketing early. Otherwise, it takes too long to learn from your users/iterate on the product.

[0] https://telointerview.com/

I used Laravel Spark for telointerview.com, I coded the base app in React so I use iframes for the spark pages. Couldn't recommend it enough, just dont be afraid to edit the vue files.

I wrote a comparison for Australian companies recently:

https://medium.com/@grmcameron/stripe-vs-braintree-for-your-...

1. Feature Related Main Takeaways:

Stripe vs Braintree feature comparison: Braintree allows payments through Paypal (This can increase conversion if you believe your customer base feels more comfortable when using Paypal, a name they will likely recognize, as their card details are not revealed.) Stripe accepts American Express while Braintree does not Braintree allows you to bring your own merchant account The stripe ecosystem is (arguably) more developed (i.e. Atlas)

2. Cost Related Takeaways:

If you are going to be bringing in under $200k revenue, then Braintree will likely be the better option in terms of cost considering their first $50k fee-less revenue. However, if you are to be earning more than $200k, then Stripe will likely be the better option in terms of cost considering their lower % fee on transactions. Additionally, the number should be re-calculated based on your international/domestic split, your estimated no. of charge backs and your average transaction size.