HN user

abought

37 karma
Posts1
Comments24
View on HN

I overlapped a bit with Chad in 2015, as he was navigating a professional transition. I wasn't in an especially high role back then- just a guy in the back of the room.

In the times I saw him since, I consistently saw someone who thought hard every day about how to help others, and didn't lose sight of the human element. Sentry worked hard to create a viable business, without losing sight of open source goals. (you can see some of his efforts at https://blog.sentry.io/authors/chad-whitacre/ )

I tell my younger colleagues to do the best work they can sustainably do... but too often in this field, the big roles become too intense to be sustained forever. I hope his new role shows him the same warmth and support that he tried to put out there for others.

Per your comment about the effect on socialization: hang in there! That space is improving!

I've never gone for drinking, and it was deeply ingrained in the social and startup experience when I was younger. This could be a bit alienating at times (it's not fun being the only sober person at a party, and there's just no right way to refuse a kegstand invite from your boss).

One of the surprising outcomes of dry january and general aging... in some ways, I'm getting to know my long-time peers for the first time. More mixer events are providing non-alcoholic drinks as a first class item, and when we socialize, instead of pressuring me to do shots, we're sharing all the other stuff I enjoyed doing all along: trading favorite tea blends, day hikes, game nights, etc.

I've never met the author, but for anyone else considering this: Welcome, new old friend!

Conferences can be truly wonderful, but not a universal replacement for publishing.

If you think journals are expensive, try sending your whole lab to a conference in another country. That may not let you in. Where some of the attendees have to fill out paperwork before talking to a foreign national. (does that ever make for awkward small talk...)

For all their many faults, journals provide access to a really wide audience, and- in theory- make it possible to form connections who wouldn't be able to meet directly.

This is a fine example of where someone's understanding of the problem runs ahead of their understanding of the solution.

A few scattered thoughts:

1. There is a difference between pre and post publication peer review. These discussions almost invariably conflate the two, but part of the runaway success of spam journals is that the benefits of pre greatly outweigh the risks of post. Historically, there was some link: if an article had problems, you would open the table of contents n months later and (might) see a letter or further discussion. Now, the table of contents is google, and many readers have weaker links to the same venue over time for followup. At the metrics level, the reputational hit of bad articles is weaker. (studies have shown that retractions are often cited with the original intent years after a correction was published)

2. The phrase "for profit" is doing a lot of work in this article. Some mega publishers, like ACS, are technically non profit member societies stapled to a mega-publisher, and have been strongly opposed to OA policies in the past. [1] https://en.wikipedia.org/wiki/American_Chemical_Society#Cont... [2] https://www.acs.org/content/dam/acsorg/about/aboutacs/financ...

3. Outsourcing trust to someone who isn't the current evil... will only get you so far. No matter who takes over publishing, scientists are going to need to evolve new ways of evaluating work and each other, as the field grows far beyond what a small network can handle. Journals are a bad metric, but how does your dean evaluate 50 people hired to be the world's leading experts on (new and emerging field)? I've read plenty of these publisher=bad screeds, and most stop there. PubPeer exists for some, Twitter walkthroughs of papers were a great thing for a while, or there's also talk of overlay journals that decouple the act of publication (as a preprint) from the review-and-prestige piece.

4. The current system does two things: (a) provides a record of work done by students, who may labor under graduation requirements to publish something, whether their project is successful or not, (b) a shared record of current state of human knowledge, be it from researchers at a small college, or google, or pharma. Goal (a) puts a lot of pressure on peer review in "low tier" journals that even the reviewers don't like to cite, and I've had my doubts as to whether this is the best yield for effort.

At various points in my career, I've had to oversee people creating data export features for research-focused apps. Eventually, I instituted a very simple rule:

As part of code review, the developer of the feature must be able to roundtrip export -> import a realistic test dataset using the same program and workflow that they expect a consumer of the data to use. They have up to one business day to accomplish this task, and are allowed to ask an end user for help. If they don't meet that goal, the PR is sent back to the developer.

What's fascinating about the exercise is that I've bounced as many "clever" hand-rolled CSV exporters (due to edge cases) as other more advanced file formats (due to total incompatibility with every COTS consuming program). All without having to say a word of judgment.

Data export is often a task anchored by humans at one end. Sometimes those humans can work with a better alternative, and it's always worth asking!

Recently, I attended an hour long meetup from a high-level AWS employee about CI pipelines using CodeCommit. Of several possible deprecation announcements, the latest of those was dated the day of the talk. (!!)

In all fairness to the speaker, except for having one of the most prominent icons across his slide deck deprecated in real time, it would have been a pretty decent talk. He even made an effort to promote the GitHub integrations as a path forward, and provide some guidance on current tooling. It was clear CodeCommit wasn't the path of most momentum, even if the degree was unclear.

A lot of the audience fragmented for various reasons, and even within the same discipline, not everyone has converged in a new place yet. I'm told mastodon got more CS/phys science people, and Bluesky got more social scientists. That's a shifting landscape though.

There was one obvious place to look for these discussions; now there are many. Changes to search tools and API access didn't help discoverability either.

Some of the departures are for practical reasons: Twitter regularly changes the rules around logged-in viewing, direct links, and promotion/ordering of posts in ways that create friction for people trying to engage in public outreach. ("this method worked yesterday" should refer to the software, not the communication, thank you very much!)

It depends. Some of the foundation-funded positions are competitive, and a few centers have surprisingly professional leadership. People are trying to organize, and- even if it's an uphill climb- there's been some improvement at the edges.

Anecdotally, some of the RSE leads I've spoken to are seeing more long-term demand than they predicted, which might lead to more room for senior roles. Currently quite a few teams (outside the big centers) seem to be priced way too low, usually explained as because they're testing the waters.... so "cheap student labor" and "one off project" is what they can afford.

Minor heretical aside: one thing I miss about old twitter is that academia was developing a real "second layer" on top of journals, where things like reproducibility could be discussed publicly. PubPeer is a partial solution, as are GitHub issues... if enough gatekeepy people really see value in code quality, norms will shift with or without mandates.

Some departments are starting to hire programmers. There's an effort to define this as a (broad) job category under the heading "Research Software Engineer":

https://society-rse.org/ https://us-rse.org/

Institutional support varies widely; some projects or teams are rather well funded for big projects and senior talent, while at other schools, the cost structure is more aimed at "one off" projects staffed by more recent graduates.

A recent grant is trying to fund this work at several schools with a history of well organized services: https://www.schmidtfutures.org/our-work-old/virtual-institut...

At networking mixers I've attended, often there are a limited number of drink tickets per person and a set event duration. But the drink tickets only covered alcohol, not other beverages. In some cases, other beverages just weren't an option at all. ("I guess you could find a water fountain? But why?")

Also, one of those recurring events was hosted at a startup that was, to be frank, known for "worse than college" level planning all around...

Eventually events started getting better at providing options. Probably a mix of several factors: the move to another city (local culture), career growth/level of people around me, and changing social patterns. (increasing interest in non alcoholic options/ more people willing to speak up)

Strange but true: I've been to a number of professional happy hours that offered free alcohol, but didn't provide other beverages. It got to the point that I started bringing my own water bottle to networking events, just in case.

I'm a big fan of providing other beverage types. Being able to sip a soda etc from the same kind of container as everyone else goes a long way towards blending in.

Location: Michigan

Remote: Open to hybrid/remote

Willing to relocate: No, but travel ok.

Technologies: Fullstack web + DevOps. Data analysis and workflows. Python. Flask, Django, Django REST framework. Celery, RabbitMQ. Terraform, Packer, Docker, Amazon Web Services (AWS). SQL (MySQL, PostgreSQL, AWS RDS, SQLite). NoSQL (MongoDB). Vanilla JavaScript/ES6 and frameworks including JQuery, D3, Ember.js, Vue.js. PyData stack (Numpy, Scipy, Matplotlib, Pandas, Jupyter), Snakemake, Nextflow, some R, MATLAB (scripting and GUI development). Some Unix/Linux administration. Unit testing and TDD with common frameworks and CI platforms (Pytest, xUnit, Mocha, GitHub actions, etc).

Résumé/CV: https://www.linkedin.com/in/abought/ (PDF on request) GitHub: https://github.com/abought

Email: abought+hnhiring [-@-] gmail [-.-] com

--

Synopsis: Research software engineer and full stack web developer, with experience in a variety of startup and R&D environments (academia, industry, nonprofit). I specialize in building collaborative tools that enable teams to create and understand large, complex datasets. I have lead or made major contributions to a number of widely used tools across multiple R&D fields of study, including features for data harmonization, visualization, and sharing. Most recently I am shifting more of my focus to the backend and working on the next generation of a tool for running a complex data annotation pipeline for large datasets in AWS. I've contributed to every part of the stack, plus some work with DevOps tools. I also maintain human subjects research certifications and have spearheaded successful federal security processes for public tools.

I work hard to help other members of my team be productive and drive continuous improvement. I'm pragmatic and flexible in choice of technologies, and like working with smart, caring people to solve problems bigger than one person can do alone. There's always something new to learn!

"Anyone in the training dataset"?

A big unanswered question in the age of AI: how does a system of law work when breaking one law is bad, but the product of breaking many laws is totally exempt?

We're starting to see the milder form of this in debates around authorship and copyright. But when your AI model requires a shockingly large quantity of clearly verboten material as input, what is one to make of the output?

I've also had luck with grocery stores. Especially the produce boxes: they have thick walls and are good for moving dishes.

From their side, they are saving the hassle of breaking down and disposing boxes. So for best results, ask them the preferred pick up time and stick to it. (there's a fine line between "less work disposing of boxes" and "tripping over stuff taking up space all day". If people feel appreciated, they can be remarkably kind.)

Oftentimes, deliveries come on specific days- for any store you ask, plan ahead slightly to work with their schedule. Not every store has the room or space to help, but some will try if they can. (Trader Joe's, for example, tends not to waste space on frivolities like storage. Or parking.)

Moving to a new place is a lot! I hope you find good friends and neighbors when you get there. Taking time to build roots before the new job gets hectic can make a big difference to quality of life later.

Docusaurus 2 Beta 5 years ago

This looks interesting!

In the past, I've been a big fan of automatic documentation generators (jsdoc, openapi, etc), because keeping a markdown file full of function names and arguments up to date by hand was painful- but I don't like that those systems have little room for prose content like guides or tutorials.

Does Docusaurus support both types of information? The examples I've browsed so far seem to involve hand-edited API docs (example: Babel - https://raw.githubusercontent.com/babel/website/main/docs/pa...). I'd love to see a system that supported building and showing API docs and prose guides in one site, or at least allowed automated cross-linking in a way that could be kept up to date.

I'd concur: the university is the wrong unit-of-ban.

For example: what happens when the students graduate- does the ban follow them to any potential employers? Or if the professor leaves for another university to continue this research?

Does the ban stay with UMN, even after everyone involved left? Or does it follow the researcher(s) to a new university, even if the new employer had no responsibility for them?

For web development, error monitoring services like Sentry are a breath of fresh air: instantly find the exact lines of code that are causing problems, even in front end code, before anyone files a bug report.

Any tool that automatically captures a lot of data needs to be used with care (eg verifying that no sensitive data is sent to the server), but many tools in this space make an effort to scrub the most common fields. In practice, the payoff has been worth it most of the time.

Per the budget spreadsheet linked elsewhere, preprints seem to be 20+% of site traffic, but only 0.4-1% of storage costs for the overall site.

This could indicate that many preprints are posted as standalone artifacts (eg not every preprint would include rich datasets alongside the PDF). In effect, people could be skipping the workflow and just using it for the preprint.

For bandwidth usage, they're estimating 20-25% of site traffic. Hosting costs are not a trivial expense, but they're not the main driver of the budget, and the bigger uncertainty seems to be on labor costs.

I think the best way to understand the numbers would be to look at the breakdown of costs. Two links that were shared from the twitter thread discussing this might be useful:

Preprints cost (projected vs actuals): https://docs.google.com/spreadsheets/d/1V0vKrf50K667CqM3e4S2...

Org finances: https://cos.io/about/our-finances/

A few things stand out: 1. Preprints are pretty new. You're not just hosting PDFs on an s3 bucket in maintenance mode- you're also wrangling authors with very different workflows to use your platform. This means building tools for moderation or retraction, and long handholding to recruit 26 partner groups, some of them started as grassroots efforts without their own institutional history. Each group may have their own ideas about governance. (each research field may do things in different ways)

2. In that light, the projected personnel costs are.... not high. The spreadsheet claims that 22% of page traffic going to preprints, but the original 2019 forecast called for ~$7k budget on developers + QA, total. At market rates, that's... a small fraction of the annual cost of a single developer? (their team page lists 10-15 devs on staff)

3. Compared to the overall organizational finances, it suggests that if anything, some of the cost of running the service is being spread across their other offerings. The original vs modified forecasts for 2019 seem a bit, well, different- it's likely that the costs are still being worked out, and may be dependent on hard-to-predict growth.

It's also very notable that this hubbub seems to involve a relatively small amount of cash: the proposed funding model is a 60-40 split, with the service share divided among up to 26 groups. That says a lot about the role of building institutions to support preprints long term, and the need to help grassroots initiatives mature if we want to keep these services active.

https://cos.io/our-products/osf-preprints/ "...this fee structure accounts for $87,976 in contributions by the preprint services toward maintenance costs (38% of total)"

Disclaimer: I don't speak for any of the groups involved in this process, and comments are based only on these public documents. There may well be other numbers or context.