This worked well for me ... I pointed claude at https://agentgrade.com/skills.json and let it incrementally increase its score.
HN user
itsderek23
https://github.com/itsderek23/
How I'm using git/Github has changed with agentic coding. However, I'm not using swarms of agents to write code, so it's bit hard for me to decipher the JTBD of gitbutler.
Another take I've seen is https://agentrepo.com/, which is light-weighted hosted git that's easy for agents to use (no accounts, no API keys, public repos are free). There are large parts of the GitHub experience I'm no longer using (mostly driving from Claude), so I think this is an interesting take.
Nice work! I'm working on a similar standalone DevOps AI Agent (OpsTower.ai). This post shows how the agent is structured and how it performs against a 40 question evaluation dataset: https://www.opstower.ai/2023-evaluating-ai-agents/
Thanks!
Did you hit token limits?
While i used TikToken to limit the message history (and keep below the token limit), generally I found that I didn't get better completions by putting a lot of data into the context. Usually the completions got more confusing. I put a limited amount of info into the context and have generally stayed below the token limit.
Are you storing message/ chat histories between sessions
Right now, yes. It's pretty important to store everything (each request / response) to debug issues with prompt, context, and the agent call loop.
+1. Jonathan is great to work with if you are in a similar position as Baremetrics.
Great Caleb - makes sense. Thanks!
This certainly looks like a cleaner way to deploy an ML model than SageMaker. Couple of questions:
* Is this really for more intensive model inference applications that need a cluster? It feels like for a lot of my models, a cluster is overkill.
* A lot of the ML deployment (Cortex, SageMaker, etc) don't see to rely on first pushing changes to version control, then deploying from there. Is there any reason for this? I can't come up for a reason why this shouldn't be the default. For example, this is how Heroku works for web apps (and this is a web app at the end of the day).
Scout also detects these for Django, ordering by the most performing N+1s: http://blog.scoutapp.com/articles/2018/04/30/finding-and-fix...
A lightweight approach we've started at my company:
1. Create a GitHub Repo dedicated to user-facing issues (https://github.com/scoutapp/roadmap)
2. Customers can subscribe to issues they are interested.
3. When resolving an issue, we reference it in the git commit, which closes the issue and notifies the issue subscribers.
We're a developer tool, so it's a familar flow for our customers.
Hi Derek! You've helped me with my employer's Scout configuration in Slack :)
Small world!
Is there an automated way of getting the average of a performance metric (eg Time spent in AR) over N requests?
I'm assuming you mean w/Chrome dev tools + server timing?
Not that I'm aware of...DevTools is an area I'd like to explore more though.
Author here.
The server timing metrics here are actually extracted from an APM tracing tool (Scout).
Tracing services generally do not give immediate feedback on the timing breakdown of a web request. At worst, the metrics are heavily aggregated. At best, you'll need to wait a couple of minutes for a trace.
The Server Timing API (which is how this works) give immediate performance information, shortening the feedback loop and allowing you to do a quick gut-check on a slow request before jumping to your tracing tool.
but I think it's an API limitation
Author here - I believe that's the case. There isn't a way to specific start & end time: https://w3c.github.io/server-timing/#dom-performanceserverti...
That said, the spec also mentions:
To minimize the HTTP overhead the provided names and descriptions should be kept as short as possible - e.g. use abbreviations and omit optional values where possible.
I could see significant issues if we tried to send data in timeline fashion (such as creating a metric for each database record call in an N+1 scenario).
One idea: pass down an URI (ie - https://scoutapp.com/r/ID) that when clicked, provides full trace information.
Author here.
Application instrumentation - whether via Prometheus, StatsD, Scout, New Relic - solves a very different problem than this. The server timing metrics here are actually extracted from an APM tool (Scout), so you get the best of both worlds.
With those tools, you do not get immediate feedback on the timing breakdown of a web request. At worst, the metrics are heavily aggregated. At best, you'll need to wait a couple of minutes for a trace.
Profiling tools that give immediate feedback on server-side production performance have their place, just like those that collect and aggregate metrics over time.
In my experience, it's very difficult to tie profiling data from generic profilers to specific requests, then to the specific lines-of-code triggering the problems.
This is important because many performance conditions don't reveal themselves all of the time: for example, it's very common that an issue might only be a problem for your largest customers. The context is really important.
Scout has a production-safe profiler for Ruby apps that builds on the wonderful StackProf gem that does this: http://help.apm.scoutapp.com/#scoutprof
Rust
NodeJS
Go
The best approach for me is finding a shared connection that will do an intro:
* It validates you * It puts the other person "on the hook". Not good form to completely ignore an introduction.
Just Rails right now - Sinatra is high on the list. We're testing it for our apps internally now.
Scout Founder (Derek) here. If you aren't running Rails, signup here and we'll email you when we support your language/framwork: https://apm.scoutapp.com/beta_invites/new.
Our source for the charts is here: https://github.com/scoutapp/scout_realtime/blob/master/lib/s...
It's not yet in a state for plug-and-play usage in other projects. If you're looking to rollout smooth-scrolling charts quickly, checkout http://smoothiecharts.org/.
Good points. Thanks!
Ah thanks - I haven't seen this yet my laptop. We'll keep an eye on it.
Yeah - we thought about this, but decided to get started in Ruby since it's the fastest way for us to ship. Go is definitely interesting.
Part of the scout_realtime team here...swap is important. Displaying it the future is possible.
In fact, fire up the console on the project homepage and type "metrics.memory". We're capturing it, just not displaying it yet on the screen.
Sorry. Just FYI, we've clocked the CPU usage of the scout_realtime daemon at 1% on an Intel Xeon 2.40GHz CPU.
Sorry - no FreeBSD support yet.
We've clocked the CPU usage of the scout_realtime daemon at 1% on an Intel Xeon 2.40GHz CPU. Memory usage is around 22 MB. If you turn off the metric collection (by clicking the pause button on the web page), CPU usage will effectively drop to 0%, and you'll still be able to visit the web page and re-enable metrics at any time.
Thanks again for reporting - we've released version 1.0.1 to fix the issue:
gem install scout_realtime
Note that OSX support is limited as there is no "/proc" support.
Thanks - I can reproduce with that output. Opened an issue on github - we'll fix: