We also use proxies with CodeRabbit’s sandboxes. Instead of using tool calls, we’ve been using LLM-generated CLI and curl commands to interact with external services like GitHub and Linear.
HN user
gillh
www.coderabbit.ai
This topic has been done to death. Too many OSS options in the last ~10 years with very little differentiation.
Let's talk LLMs instead.
Have to provide precise instructions to LLMs to get anything useful.
Instructions for operating the "Holy hand grenade of Antioch" from Monty Python and the Holy Grail are a good example:
''First shalt thou take out the Holy Pin. Then shalt thou count to three, no more, no less. Three shall be the number thou shalt count, and the number of the counting shall be three. Four shalt thou not count, neither count thou two, excepting that thou then proceed to three. Five is right out. Once the number three, being the third number, be reached, then lobbest thou thy Holy Hand Grenade of Antioch towards thy foe, who, being naughty in My sight, shall snuff it.'
Combination of zsh completions with fzf picker and gh copilot for CLIs.
I have a good setup here that you can learn from - https://github.com/fluxninja/dotfiles
Anyone interested in load shedding and graceful degradation with request prioritization should check out the Aperture OSS project.
When people ask us about any other AI tool, our standard reply is - "Please try both tools to see the difference and choose the one you like."
We will appreciate the same courtesy from our competitors.
Cheers!
I see that the handful of Ellipsis "buddies" here are upvoting your post. :)
CodeRabbit employees wouldn't usually be commenting here to spoil your "moment," but this reply is entirely wrong on so many levels. The fact is that CR is much further along the traction (several hundred paying customers and thousands of GitHub app installs) and product quality. Most of the CR clones are just copying the CR UX (and it's OSS prompts) at this point, including Ellipsis. The chat feature at CR is also pretty advanced - it even comes with a sandbox environment to execute AI-generated shell commands that help it deep dive into the codebase.
Again, I am sorry that we had to push back on this reply; we usually don't respond to competitors - but this statement was plain wrong, so we had to flag it.
Anyone looking to build a practical solution that involves weighted-fair queueing for request prioritization and load shedding should check out - https://github.com/fluxninja/aperture
The overload problem is quite common in generative AI apps, necessitating a sophisticated approach. Even when using external models (e.g. by OpenAI), the developers have to deal with overloads in the form of service rate limits imposed by those providers. Here is a blog post that shares how Aperture helps manage OpenAI gpt-4 overload with WFQ scheduling - https://blog.fluxninja.com/blog/coderabbit-openai-rate-limit...
Interesting use-case: We recently started using ast-grep at CodeRabbit[0] to review pull request changes with AI.
We use gpt-4 to generate ast-grep patterns to deep-dive and verify pull-request integrity. We just rolled this feature out 3 days back and are getting excellent results!
Comments such as these are powered by AI-generate ast-grep queries: https://github.com/amorphie/contract/pull/100#discussion_r14...
FluxNinja [0] founder here. I developed an in-house AI-based code review tool [1] that CodeRabbit is now commercializing [2].
I did it because of the increasing frustration due to the time-consuming, manual code review process. We tried several techniques to improve velocity - e.g., stacked pull requests, but the AI tool helped the most.
I work at FluxNinja
CodeRabbit is a customer of ours - https://docs.fluxninja.com/blog/coderabbit-openai-rate-limit...
In addition to code generation, our team has found the new AI code review tools to be quite useful as well. We use CodeRabbit and we keep finding issues/improvements in every other PR.
We haven’t used functions as our text parsing has been pretty high fidelity so far. It’s because we provide an example of the format that we expect. We didn’t feel like fighting too hard with LLMs to get structured output. You will also notice that our input format is not structured as well. Instead of unidiff format we provide the AI side-by-side diff with line number annotations so the it can comment accurately - this is similar to how humans want to look at diffs.
Our OSS code is far behind our proprietary version. We have a lot more going on over there and we don’t use functions in that version as well.
We have been doing this at CodeRabbit[0] for incrementally reviewing PRs and allowing conversations in the context of code changes, giving the impression that the bot has much more context than it has. It's one of the few tricks we use to scale the AI to code review even large PRs (100+ files).
For each commit, we summarize diff for each file. Then, we create a summary of summaries, which is incrementally updated as further commits are made on a pull request. This summary of summaries is saved, hidden inside a comment on a pull request, and is used while reviewing each file and answering the user's queries.
Some of our code is in the open source. Here is the link to the relevant prompt for recursive summarization - https://github.com/coderabbitai/ai-pr-reviewer/blob/main/src...
[0]: coderabbit.ai
Prioritized load shedding works well as a last resort [0]. The idea is simple -
- Detect overload/congestion build-up at the database
- Apply queueing at the gateway service and schedule requests based on their priority
- Shed excess requests after a timeout
[0] https://docs.fluxninja.com/blog/protecting-postgresql-with-a...
Fascinating historical insight!
We have a team of 20 engineers currently working on solving this problem in the context of API requests and service chains. Do you know JMS @ Penn? Asking because he did some work in ATM networks, QoS etc. He is advising us on the project.
This is the link to the project: https://github.com/fluxninja/aperture
We built a weighted-fair queueing scheduler as well - https://docs.fluxninja.com/concepts/scheduler/load-scheduler
We built a fair scheduler for APIs and it's in the open source - https://docs.fluxninja.com/concepts/scheduler/load-scheduler
I wish more people knew about this project!
You should check out the Aperture flow control system - https://github.com/fluxninja/aperture
We built a weighted fair queueing system in the Aperture flow control system to help alleviate the API load pressure - https://docs.fluxninja.com/concepts/scheduler/load-scheduler
And in addition, we are investing in the graceful-js library to handle 429 and 523 codes returned by the Aperture system - https://github.com/fluxninja/graceful-js
Very interesting blog post! Our team has been working intensively in this area for the last couple of years - flow control, load shedding, controllability (PID control), and so on.
We have open-sourced our work at - https://github.com/fluxninja/aperture
We also did a Twitter Spaces discussion with Kelsey earlier today - https://twitter.com/kelseyhightower/status/16893552848026296...
We would love feedback from folks reading this blog post!
Disclaimer: I am one of the co-authors of the Aperture project. There are several interesting ideas we have built into this project, and I will be happy to dive into the technical details as well.
Very interesting blog post! Our team has been working intensively in this area for the last couple of years - flow control, load shedding, controllability (PID control), and so on.
We have open-sourced our work at - https://github.com/fluxninja/aperture
We would love feedback from folks reading this blog post!
Disclaimer: I am one of the co-authors of the Aperture project. There are several interesting ideas we have built into this project and I will be happy to dive into the technical details as well.
AI helps as well. For instance, our company has extensively used coderabbit.ai since its inception.
Came across this very interesting GitHub pull request of an AI bot reviewing another AI bot’s generated code!
Is this the way of the future?
Aperture project's policy language is indeed inspired by Control Systems, especially PID Control.
Thanks for the feedback - This blog could have been better and less aggressively distributed. At the same time, we should consider the credible body of work behind this. It's not AI-generated spam but backed by actual code that provides a reference implementation behind the idea.
It's a blog on an open-source project that precisely tells you how to implement adaptive rate limiting.
Just click around a bit:
- https://github.com/fluxninja/aperture
- https://docs.fluxninja.com/use-cases/adaptive-service-protec...
Note: I am one of the authors' of this project.
While the tech these companies built (stateful reconstruction) is quite hard to engineer, large-scale observability platforms like Datadog were much more successful as they could cover a much larger surface area with minimal incremental effort. Even APM companies like New Relic and AppD struggled to justify the value they were delivering with their well-engineered code-level agents in the face of large-scale server/process metrics collection done by the Datadog agent.
In addition, these techniques have been prone to event misses (partial reconstruction) in case the ring buffer overflows and are ineffective in the face of SSL traffic.
Companies that pioneered the stateful reconstruction of APIs from raw network events -
- Netsil [0] (acquired by Nutanix) - Academic spin-out from UPenn
- Pixie Labs [1] (acquired by New Relic)
[0] https://thenewstack.io/netsil-visualizes-performance-microse...
[1] https://techcrunch.com/2020/12/10/new-relic-acquires-kuberne...
Good observation. We were able to prevent issues due to limited context by sending the entire file in the prompt and the results were pretty amazing. We eventually reverted the change due to 2 reasons -
1. Limited tokens: 8K tokens is sometimes not enough to hold the entire file. Perhaps 32K tokens (or 100K Claude2 tokens) can help circumvent this.
2. Cost: GPT-4 is super expensive. Our current usage is roughly $20 a day but when we send files the usage shot up to $60 a day or so.
Our hope is that both the cost and token limit improve in the future so that we can send the entire file in each review request.
It’s all about providing relevant context in the prompts and using a better model where reasoning is needed, eg gpt-4.
If you are curious, some of our prompts are in the open source.
Here is the link to our open source version - https://github.com/coderabbitai/openai-pr-reviewer
Look at all these super knowledgeable “engineers” down voting this.
Load shedding is exactly what would prevent this in the future, scaling up capacity and adding a cache is only going to help buy some time.
Wish I could have a dollar for each time SREs duct tape rather than solve problems with systems thinking.