HN user

gerner

122 karma
Posts2
Comments50
View on HN

This is just the beginning for sharing metrics like this. We're looking forward to feedback on this batch of results which will help form other analyses that would be interesting to see about the dev process in general.

We believe there's a lot of information locked in dev tools like code review that can shed quantitative light on how those processes lead to successes or failures in development. This is in partnership with the kind of anecdototes and best practices experienced devs and managers already frequently share.

Nick here, CEO/Founder of AutoDevTech. The team and I are proud to release AutoDev Analytics: reports and benchmarks measuring dev processes based on thousands of big, popular projects on GitHub. We want to share what we’ve learned about the dev process and code review, get feedback, and open up free analytics that compare your team against thousands of other code bases.

We’ve got stats showing most code doesn’t get any feedback in code review[1], unreviewed code is twice as likely to churn out over a year than reviewed code[2], and teams spend about 3 hours per week reviewing code[3]. And we’ve got benchmarks for dozens of other metrics like code review turnaround and monthly codebase churn.

Take a look, sign up for analytics and let us know what you think.

[1] http://www.autodevtech.com/benchmark/scorecard/denoland/deno... [2] http://www.autodevtech.com/benchmark/scorecard/nodejs/node/c... [3] http://www.autodevtech.com/benchmark/scorecard/puppeteer/pup...

Would love to see a little more on the landing page than just one screenshot. How about more description of features, why this is better than comparable options paid and free. How about a live version we can see in action? What does deployment of an agent look like?

AutoDevTech | Software Engineer | Full-time | Seattle/Remote

Come join the team as one of our first engineers building tools to help developers write better code. We’ve got a dataset of code and outcomes like review feedback and bugs and we’re using that to train ML models to predict what happens to code as developers build new features.

We’re looking for versatile developers interested in an early-stage startup experience building user experiences to assist devs across the development lifecycle (think code reviews, testing, fixing bugs). There are a lot of challenges presenting information like this in web apps and 3rd party integrations. We’re looking for someone to help define and solve those challenges.

The team consists of industry vets, led by an engineering and product head with multiple startup successes from founding to exit. We’re looking to grow a supportive, challenging, fast-paced culture focused on building some of the best technology and products in the industry.

https://www.autodevtech.com/careers

Fascinating video. It's cool to see a collaboration of child development and reinforcement learning in action and to hear a research speak about both in an experimental context.

The comment in a different thread about Diamond Age comes to mind, and here we see some of those elements playing out.

I think this is a reasonable question. Usage based pricing seems like a trending topic lately and I hear Slack called out as a good example. Monthly seat licenses were a practice that existed long before more obvious "usage based pricing" models became prominent.

I'd love to hear more thoughts on the topic.

Flat Data 5 years ago

GitHub Flavored Markdown seems like a nice extension to Markdown to me. Fenced code blocks? Great idea. Lots of other flavors of Markdown do the same thing. I don't know who's the leader or follower here, but I'm glad they're doing it. I'm not sure what's the gold standard for wikis, but they all seem like the kind of thing every vendor has similar, flawed, good-enough solutions for. And I know there are other thoughts around how to manage merges, but having a merge commit (or a squash merge or a fast forward) seems like a reasonable contender for handling a feature branch. But maybe there's something I'm missing? I guess any hegemony is bad for innovation?

Are there a lot of walled gardens that only allow sign-in with GitHub? That's not really an issue I've run into. I can't think of any site I want to invite my aunt/uncle/cousin to log into that only accepts GitHub login. In fact, I'm not sure there's a lot of tools I want my colleagues to use that require a GitHub login that isn't already tied to a GitHub hosted repo.

I would love to hear what I'm missing though.

Flat Data 5 years ago

Yes. Or S3 bucket, or whatever. The thing I'm getting at is, can we use GitHub actions for application tasks like web sraping that need compute and network access, but that don't really do much with with a git repo. Does GitHub want to support that?

Flat Data 5 years ago

See the comment from @jasoncwarner about GitHub actions being a platform for much more than CI.

I wonder how far that extends to non-GitHub provided services. For instance, could we leverage GitHub actions, perhaps even Flat Data, to scrape some web site and store it (perhaps uploading elsewhere) in a more comprehensive way vs. storing some small snippet of the data in a git repo?

Flat Data 5 years ago

Agree, it's important that we keep an eye on things and, however we can, hold MSFT and GitHub accountable to keep up the good showing.

We've seen new features launched (e.g. this one) long enough after the acquisition that much (most, all?) of the work happened in the post acquisition environment that I'm optimistic. But I've been wrong before.

Flat Data 5 years ago

I don't know much about Flat Data, but I'm impressed with how much GitHub is doing as GitHub since the MSFT acquisition. They continue to offer compelling services to developers, and increasingly to enterprise customers. All without abandoning much of what made GitHub great: a focus on developers and easy to access dev productivity.

Notice the prominence of the VSCode integration here. Notice the dramatically increased presence of MSFT on GitHub in general. It seems like they've managed to integrate these two cultures and product-sets in sensible ways. Given how hard big integrations like this are to pull off, I feel like the community really dodged a bullet in terms of access to products/tools.

AutoDevTech | Developers, Data Scientists | Full-time | Seattle/Remote US

I’m building AI/ML to help devs write better code by giving customized, actionable feedback learned from existing coding patterns and outcome indicators like review feedback, bugs and production data. I’ve got funding, a very large and growing dataset, labels and a ton of technical challenges to solve.

I’m hiring the initial team: data scientists in ML/NLP, topic modelling, classification of semi-structured data, plus developers that can build UI/UX, 3rd party integrations and support large-scale data pipelines. I want people looking to learn a ton and launch a product from scratch.

We’re going to be part of a big change in how computing supports the creative process throughout software development.

Contact careers@autodevtech.com to get involved.

AutoDevTech | Developers, Data Scientists | Full-time | Seattle/Remote

We’re building AI/ML to fundamentally change the way devs write code: giving customized, actionable feedback, and researching bugs by learning from existing coding patterns.

We’ve got funding to assemble the early team and are looking for motivated, versatile developers and data scientists that want to build the initial tech/product at an early-stage startup.

Contact careers@autodevtech.com for more information.

AutoDevTech | Developers, Data Scientists | Full-time | Seattle/Remote

We’re building AI/ML to fundamentally change the way devs write code: giving customized, actionable feedback, and researching bugs by learning from existing coding patterns.

We’ve got funding to assemble the early team and are looking for motivated, versatile developers and data scientists that want to build the initial tech/product at an early-stage startup.

Contact careers@autodevtech.com for more information.

I think it's doing a little more than that:

"Caspy’s AI learns from your email history to figure out what types of emails you send replies to. With that knowledge, it will alert you only when an email needs a reply."

I also struggled with the star rating, especially in my first few days. But at some point I just got over it. After a couple of weeks using it I'm just rating stuff without too much care and it seems pretty reasonable.

I think of the Netflix scale: 1 == terrible, never show me stuff like this in this topic (I think this is actually their threshold for exclusion from a topic) 2 == on topic, but but I don't like it 3 == this is OK, but I'd rather have something else 4 == good, solid content I'd be happy reading just this stuff 5 == wow, this is great, if it comes up it better be #1

@idina_news @_b can you confirm that ratings are _only_ applied in topic? (they don't leak to other topics)

I've been in the Idina closed beta for a couple of weeks. The #1 thing I like about this is the power of "rate and refresh". The low latency feedback loop is really impressive. Maybe I don't use enough news sites, but I haven't seen anything like that before.

The breadth of content is pretty good. I was surprised I was able to build a good gaming topic that's not just AAA titles or some niche category.

I'm a little worried about rating my way into an echo chamber. Having used Idina for a few weeks, that seems like a risk. For example, 4 of top 6 stories in my gaming topic are from the same site. I really like those stories and that site, but I wonder if I'm missing stuff from sites I've never seen before.

In my experience, and I suppose depending on the data, I've found that grep is often the bottleneck for data pipeline tasks like you describe. The silver searcher (https://github.com/ggreer/the_silver_searcher) is, in my experience, about 10x faster than grep for tasks like pulling out fields from json files. It's changed my life.

pv (pipe viewer, http://www.ivarch.com/programs/pv.shtml) and top are pretty handy to measure this kind of thing. You should be able to see exactly which process is using how much CPU, and what your throughput is.

I guess I was more interested in channels that aren't visible to google analytics (i.e. anything other than PPC?)

how about banner impressions that don't result in a click and aren't run by Google? What about views on guest posts or partner sites?

If your entire marketing strategy is SEO with clickthrough + PPC + on-site content, then, yeah, seems like GA should solve the problem. But if not? Sample size of me: lots of people do marketing activities outside the Google ecosystem.

"The credit of acquiring that customer, and the cost to acquire that customer, really should be spread among all the different marketing activities."

Wow. I have been banging my head against this problem on the ad-tech-product-side for months. But this puts it really clearly. Other than having omniscient view of all marketing activities, any thoughts on even imprecisely measuring this.

You're right, quicksand all around.

However, the post illustrates an important feature I often see overlooked: the systems we build on top of are not just black boxes. We should not be afraid to really understand how they work, opening up source code, asking questions like, "what the heck is this /proc/NUMBER section of the filesystem?" And I think this post illuminates that in a pretty fun way.

True, but it seems like this is not much a disincentive compared to the revenue potential in that data. Having worked at a number of companies that have acquired data (legally) from a fraction of this many people, one should easily be able to turn those around for a comparable amount of money. I'm not saying there's $100M of revenue here (for 2 million folks). But It seems like the penalty must be much worse than the upside to make the risk adjusted expected value work out in favor of following laws and regulations.

I might agree with you if I thought the bad press Verizon is getting was actually quite costly. But I don't.

I also have an X1 Carbon, been running linux flawlessly for over a year, probably close to 2. I love it. It's great. I got some bleeding edge graphics drivers and play Strike Suit Zero without issues (but my standards are low) when I'm not having it do map reduce jobs :)

That said, we've got some Lenovo T440s in the office that come with a wifi card (Realtek RTL8192EE) that's unsupported by ubuntu at the moment (there's an active bug at ubuntu [1] that walks through the issues and has workarounds)

[1] https://bugs.launchpad.net/bugs/1239578

In some ways, this is the dark side of disrupting established industries: legal protections aren't in place yet; it's not clear what recourse people caught in never before seen corner cases have.

We (e.g. HN folks) are quick to applaud people with new ideas that are shaking things up (e.g. Airbnb, Uber, etc.) I suspect this is the sort of scenario detractors (e.g. Seattle City Council re: Uber) are worried about.

Terrible situation, clearly. Still, important to see some of these scenarios actually play out. And it raises all the right questions: what should Airbnb do going forward? what should legal protections be going forward? should someone have seen this coming? and should they be held liable for it?

(edit: typo)

Isn't Amazon S3 $0.085 / GB for the first TB/month (1/10 of what you've said)?

That's still 8.5x Google Drive. Being competitive with glacier jives with what I've heard the underlying cost of storage hardware is (still with room for profit). However, from what I hear running the request serving on top of that is fairly expensive (which is why glacier has big restrictions).

AWS says they're a cost plus org which means they drop prices as their costs drop. But that plus can be quite large.

Being #1, first to market, etc. has a big advantage. Despite knowing (or believing) these things, my org still happily spends plenty at AWS.

It's great to see some research that shows a college education is still worth it. And I certainly believe that, although I'm admittedly part of the academically trained elite.

That said, the work to "democratize post-secondary education" is still an exciting area that can co-exist with traditional higher ed. There's a whole segment of the population for whom higher ed is not appropriate (for a variety of reasons). See what the Gates foundation is doing: http://www.gatesfoundation.org/What-We-Do/US-Program/Postsec...

I'm really glad to hear someone talk about the development/debug advantages of static typing. I work in a mixed java/ruby/python shop and issues with typing come up all the time. We catch so many bugs during java compilation that we can't see in ruby/python.

Sure, if you've got great test coverage, you get that with test automation. But I can't tell you how many refactors with "reasonable" (< 100%) test coverage went flawlessly in java with just a compile/fix-compilation-errors loop. I think this is much easier than a run-tests/fix-bugs approach with a type-free language. This is especially true when you've got maturing code and lots of ownership hand-offs.

Any tips on static analysis for ruby/python out there for a curmudgeonly java dev?