HN user

andrew311

137 karma
Posts14
Comments34
View on HN

it’s difficult to look at people like Graham — people who aren’t as bright as they think they are

Graham’s (alleged) arrogance about his brightness isn’t really the issue here. Let’s face it, he is bright. That’s not what is causing this boredom/dismay, though.

The issue is that somehow the rest of us became entranced by the “cult of Graham” and his thinking about startups/founders, and collectively we made his way into the way, ostracizing those that lived their life outside the idealized startup paradigm that Graham crafted. Creation of this dismay isn’t on him alone, it’s on all of us.

Sugar Free Dark Almonds from Sees[1] has to be one of my favorite chocolate products of all time. Friends and family always gobble them up at my place so I sometimes buy them as little gifts. I normally abhor anything “sugar free” because substitutes like aspartame taste absolutely disgusting to me, but this uses maltitol, which is a sugar alcohol that achieves a more subtle sweetness with none of the gross fake sugar taste, and since I like my chocolate on the less sweet side, it strikes a perfect balance for me. Maltitol is a mild laxative, but I’m pretty sure you would have to eat multiple boxes to feel anything.

[1] https://www.sees.com/dark-chocolate/sugar-free-almonds/20037...

To me, the real issue is that content in apps is temporal with a lack of visit history, unreproducible feeds, and a lack of deep links. This means I often lose my place when a momentary switch to another app causes a refresh on the first app. Even worse, sometimes switches are accidental from a push notification popping right before I click something else at the top of my screen. Websites have this unreproducible feed problem, too, but it generally isn’t as pronounced.

A memory seared into my brain from the 90s:

I was in the highschool computer lab, working on a website for the school IT administrator. She wanted me to include a patterned background, blinking marquees, GIFs, the whole nine yards. Another student was working on the project with me, sitting next to me. We were both facing the window, away from the door. The admin left the room. I told my fellow student these choices were extremely tacky, and I questioned the life choices that led her to the point of thinking this looked good.

Turns out she came back into the room and was behind me the whole time. Boy, did I feel like a jerk. I apologized.

What a different time that was.

This argument falls into the trap of comparing the average home to the stock market. If I was primarily concerned with returns, I would never buy the average home. That would be a home in the middle of nowhere. I would only buy in major metro areas with diverse, heavily entrenched industries and a strong desire among people all over the world to live there (places like NYC). If you look at the numbers in those places, the story is very different. On home value alone, you see appreciation of 8-12% per year on average since 1990 (despite at least three recessions during that time). Covid will impact the desire to live in major metro areas but not enough to seriously impact these returns.

If I wasn't concerned primarily with returns and instead on just saving money over renting, the math is still way better in any of the top metros in the United States so long as you plan to reside there for 5+ years. And if you aren't living in a major metro, you still need to find a place to live, so you might as well make 1% annually on that money as opposed to just giving it away in the form of rent.

In my experience, renting makes sense when you need to pay for flexibility, because you aren't sure if you will stay rooted in one place for 5+ years. Otherwise, buying for many people is a win financially and has been for decades.

I should clarify that it is certainly doable for widely recognized companies, but it’s very difficult for the majority of startups that no one has heard of even if they have some success. Also, getting on these marketplaces also goes much better with company cooperation and many startups don’t have the time or willingness.

Selling private shares / options on the secondary market is near impossible if the company isn’t on something like SecondMarket. Right of First Refusal, Co-sale Agreements, and the challenges of sharing information with a 3rd party make this difficult. That said, has anyone succeeded and written about their experience?

I agree wholeheartedly. This is important and will provide balance. That said, often people have some sort of passion, their job might be a manifestation of that passion, and they want to connect with peers.

An alternative way of looking at it: whether or not face to face is important for the specifics of the job, it might be important for emotional connection between humans, fulfilling a common human desire to connect with peers and be happy on the job and _outside_ the job. In other words, we might want to be around peers for social (non-job) reasons.

This has been my experience as someone who started a fintech company and has friends in compliance departments at banks. Highly bureaucratic on both the bank and regulator side. Compliance reviews are mostly an expensive song and dance by the banks. People jump from bank to regulator and vice versa all the time so there is rampant cronyism. Combine the cost of this song and dance with the cronyism and it becomes nearly impossible for a startup without a ton of funding.

There is a major issue at stake here that should deeply concern anyone who cares about innovation in the US. I feel that HN is getting too caught up in the "scam" rhetoric and missing the bigger picture.

Whether or not Kin is a scam is an entirely different matter from whether or not Kin is a security. The SEC has jurisdiction over securities, but not over scams generally. There is a substantial debate as to whether or not Kin is a security. At the very least, a decent argument has been made that it is not. Clearly Kik and their lawyers think they can win. Personally, I largely agree with the logic laid out by Kik in their Wells Response to the SEC[1]. HN user elliekelly gave other great examples of of the difficulty in applying securities laws in cases like this[2].

We have to be very careful ceding ground to the SEC on what constitutes a security. This is absolutely worth fighting if you care about innovation in the United States. Many of the donors to Kik's Defend Crypto campaign couldn't care less about Kin. That is not the point. However, they do care immensely about the ramifications this case could have for businesses.

Businesses should be experimenting with new ways to finance companies. Perhaps that is through the sale of tokenized products or virtual currencies. Maybe they will provide great alternatives to venture capital over the long term. The label of "security" is an onerous one that creates massive obstacles to that innovation, especially for small companies. If the SEC succeeds in labeling Kin a security, what could have been the beginning of an innovative step towards new business models for fledgling startups is now forever regulated away into obscurity in the US. Meanwhile, businesses in other countries get to keep experimenting. The lackluster guidance and unpredictable enforcement has already led to companies taking their business elsewhere. It's a shame really.

If there is even a sliver of a doubt as to whether or not Kin constitutes a security, we should not be so quick to cede ground to the SEC. If we let the SEC go unchallenged, they will expand their reach, becoming more entrenched and widening the scope of what constitutes a security. Gaining ground back becomes harder over time, especially if the SEC wins court cases.

If Kin truly is a scam, we have a multitude of ways to prosecute them without involving the notion of securities. If nothing else, we always have regular contract or tort law if there were any contractual misrepresentations or intent to defraud. We don't need the label of "security" or action from the SEC for these kinds of claims, and this approach would be perfectly adequate. There are a number of government agencies that could bring these sorts of cases and fight for the public. Trying to prove it is a security at the same time is simply regulatory overreach.

Bottom line, maybe Kin is a scam (I don't think so), maybe someone should do something about it, but let's be careful about expanding the scope of what constitutes a security. There are plenty of ways to prosecute Kik without ceding that ground.

Side point: if the SEC succeeds in labeling Kin a security this creates all sorts of logical incongruities with past no-action letters or lack of enforcement in other areas. For example, if Kin is a security, why were the San Francisco Giants given a no-action letter for pre-sales of stadium seats "all of which were initially sold to fans prior to the Park’s opening day" which could be resold through "a service that would facilitate the resale of Charter and Club seat licenses"?[3] Sure, the Giants made a buyer represent that they were "not acquiring the [seat] as an investment and has no expectation of profit", but do we really think that stopped people from buying with the intent to profit? ICOs put the same representations in some of their pre-sale agreements, and we all know that did not stop people. What amount of intent to consume vs resale is appropriate? Broadway theater shows do the same sort of pre-sales of seat licenses, and we all know how much people profit from the resale of successful shows. This checks all of the boxes of the Howie Test (paid money, expectation of profit, dependency on managerial efforts). How come the SEC does not bring action there? I don't see fair and even enforcement of the law, which really brings the efficacy of it all into question.

[1] Kin Wells Response, https://www.kin.org/wells_response.pdf

[2] HN comment by elliekelly, https://news.ycombinator.com/item?id=20098363

[3] https://www.sec.gov/divisions/corpfin/cf-noaction/sfba022406...

Share grants would be seen as income by the IRS and most states and taxed at their Fair Market Value. Options on the other hand usually qualify as Incentive Stock Options that aren’t taxed at grant time and “when exercised, it isn't necessary to pay ordinary income tax. Instead, the options are taxed at a capital gains rate.” [1]

Options are better up front because there is no outlay for the employee. They are a hassle down the road. However, if you exercise during a liquidation event your tax liability is probably covered.

Stock is a pain upfront unless granted before the first round of funding or any real revenue when the stock value is very little. They are easier down the road, though.

Just my two cents. HackerNews, please correct any errors in logic or how this stuff works.

1. https://www.investopedia.com/terms/i/iso.asp

Correct, the analysis comes with the disclaimer that I am not a lawyer, and it is not legal advice. I am active in the space and have consulted leading lawyers for my own activities, so I come with some knowledge, but obviously this is a quickly evolving space.

For anyone buying tokens and considering SAFT offerings, this is a good read from a securities lawyer, albeit an English one (not US):

https://prestonbyrne.com/2017/08/04/thoughts-on-the-saft/

This is also a great post in terms of highlighting pitfalls and challenges:

https://medium.com/@twobitidiot/losing-alpha-why-most-new-cr...

Good question. Here's one benefit. With Kinesis you can batch a bunch of writes to S3 that would have otherwise resulted in many small files in S3.

In other words, you can make small writes to Kinesis and then read out in larger amounts and write larger files to S3. This is a huge optimization for any job that runs across the data in S3. Many small files can really undermine performance in something like Hadoop MapReduce because of the additional request overhead.

Yes, good point. This would provide effectively "infinite" backing storage. There might be some hurdles to overcome, though. For example, when you delete a file, will EBS know that the blocks are now free and thus can be decommissioned? This might mean the whole stack needs to support things like TRIM. I'm not sure the rest of the stack is smart enough yet. I'd love to hear from a storage/FS expert on this.

Edit: coincidentally, I just saw this article about XFS which observes the following:

"Over the next five or more years, XFS needs to have better integration with the block devices it sits on top of. Information needs to pass back and forth between XFS and the block device, [says Dave Chinner]. That will allow better support of thin provisioning." https://lwn.net/Articles/638546/

EFS is a great addition to AWS. We have SAN as a service via EBS, now we get NFS as a service. Great.

The question (for me) now becomes "where do we go from here?"

Infinite NFS is great, but what I've always wanted is infinite EBS that is fully integrated from file system to SAN. In other words, something that behaves like a local file system (without the gotchas of NFS like a lack of delete on close), but I don't have to snapshot and create new volumes and issue file system expansion commands to grow a volume. I want seamless and automatic growth.

Furthermore, there's so much local SSD just sitting around when using EBS. I want to make full use of local SSD inside of an EC2 instance to do write-back or write-through caching. I could do this in software, but maybe there's an abstraction begging to be made at the service level.

Throw in things like snapshots, and this would make for a fairly powerful solution, and it would certainly remove a lot of operational concerns around growing database nodes and such.

Don't get me wrong, you can pull together a few things and write some automation to do this today. You could use LVM to stitch together many EBS volumes, add in caching middleware (dm-cache, flashcache, etc.), and then automate the addition of volumes and file system growth. However, it's clunky, and there's an opportunity to make this much easier.

I recognize that what I'm describing doesn't serve the same purpose as NFS - for example, EBS isn't mountable in multiple locations at once - but I'd really like to see the "seamless infinite storage" idea applied to EBS.

Aside from asking about internal company details such as financing valuations, you should also do research on similar companies that are a bit more mature. Find out details on their financings, revenue, exits, etc. It'll give you some idea of what could happen.

Question about Composite Hash Keys that someone might have the answer to (or be able to relate to other known implementations):

The composite key has two attributes, a “hash attribute” and a “range attribute.” You can do a range query within records that have the same hash attribute.

It would obviously be untenable if they spread records with the same hash attribute across many servers. You'd have a scatter-gather issue. Range queries would need to query all servers and pop the min from each until it's done, and that significantly taxes network IO.

This implies that they try to keep the number of servers that host records for the same hash attribute to a minimum. Consequently, if you store too many documents with the same hash attribute, wouldn't you overburden that subset of servers, and thus see performance degradation?

Azure has similar functionality for their table service, requiring a partition key, and they explicitly say that there is a throughput limit for records with the same partition key. I haven't seen similar language from Amazon.

Whether you scatter-gather or try to cluster values to a small set of servers, you'll eventually degrade in performance. Does anyone have insight into Amazon's implementation?

This is true. If the assumption is that it's a private branch, then other people shouldn't care if you push -f because no one else should be using it.

Sometimes there are cases where people want to pull a private branch because they are working on something that is in the same code path but will be deployed after the private branch is integrated and deployed. They want to work off the newest code and avoid a larger merge to their private branch later. Would rebasing that private branch make their life harder? If so, one could always stage changes in a feature branch at stable points for them. Thoughts?

Basically, my understanding is that push -f can be a hassle for others to pull if they made commits to the same branch already. You're totally right that if it's truly a private branch, though, this should be irrelevant.

I'm wondering how people address one of the scenarios raised in the post, specifically this:

"It’s safest to keep private branches local. If you do need to push one, maybe to synchronize your work and home computers, tell your teammates that the branch you pushed is private so they don’t base work off of it.

You should never merge a private branch directly into a public branch with a vanilla merge. First, clean up your branch with tools like reset, rebase, squash merges, and commit amending."

I'm wonder how people address cleaning a private branch that has been pushed (when your goal is to get its changes into master cleanly). Rebasing the private branch is pretty much out of the picture since it has been pushed (unless you don't care about pushing it again). I can see some ways of doing this:

1) You could do a diff patch and apply it master, then commit.

2) You could checkout your private feature branch, do a git reset to master in such a way that your index is still from the private, then commit it. Ex:

currently on private branch git reset --soft master

Now all the changes from the private branch are changes to be committed on master. This is easy, but it puts everything in one commit.

If you wanted to do a few commits for different, but stable points, but you already pushed the private branch and can't rebase it, you could instead do "git reset --soft" on successive points in the private branch commit chain, committing to master as you go.

If you wanted to reorder commits from the private branch, I guess you could rebase the private branch (which means you can't push again since you pushed it already), then do the tactic from the last paragraph, then ditch the private branch cause it's no longer pushable.

Does anyone have better ways of putting changes to master for private branches that have already been pushed?

Michael, interesting presentation.

I'm the guy that created the "Ops per second" slide included in the presentation. It was originally from my presentation at MongoNYC:

http://www.10gen.com/presentation/mongonyc-2011/optimizing-m...

I almost regret including that chart because a lot of people are citing it out of context without explanation. That chart illustrates a very extreme case, and it makes an exaggerated point so that people know to be conscious of working set size vs RAM.

FooBarWidget made some very valid points about working set elsewhere in these comments. It's important to note that any database will suffer when data spills out of RAM. The degree to which it suffers varies in different workloads, DBs, and hardware.

Databases like PostgreSQL do better in some cases because they have more granular/robust locking and yielding, so disk access is handled better and generally leads to less degradation in performance.

For example, in MongoDB, when you do a write, the whole DB locks. Additionally, it does not yield the lock while reading a page from disk in the case of an update (this is changing). So if you spill out of RAM, and you update something, now you need to traverse a b-tree, pull in data, and write to a data extent, possibly all against disk, while blocking other operations.

In other databases, they might yield while the data is being fetched from disk for update (or have more granular locks), so a read on an entirely different piece of data can go through. That other piece of data might reside on another disk (imagine RAID setups), so it is able to finish while the write is also going. Beyond that, MVCC DBs can do a read with no read lock at all.

Yes, MongoDB lacks this robustness, and it's why it suffers worse under disk bound workloads than other databases. However, 10gen is very aware of this, and they are making a strong effort to introduce better yielding and locking over the next year.

Regarding indexing and working sets, FooBarWidget points out that there are very naive ways to approach indexing that lead to a much larger working set than is necessary. With very simple tweaks you can have a working set that only includes a sub-tree of the index and smaller portions of data extents. These tweaks apply to many DBs out there.

The classic example is using random hashes or GUIDs for primary ids. If you do this, any new record will end up somewhere entirely arbitrary in the index. This means the whole index is your working set, because placing the new record could traverse any arbitrary path. Imagine instead that you prefix your GUID with time (or use a GUID that has time as the first component). Now your tree will arrange itself according to time, and you will only need the sub-trees for the data that has time ranges you frequently look at. Most workloads only reference recent data, so this helps a ton. You’ll still find yourself doing stuff like this elsewhere in other databases to improve locality.

What’s really interesting is when you look at MVCC, WAL log systems and how they compare. They get a huge win by not needing locks to read things in a committed state. They also convert a lot of ops that involve random IO into sequential IO, but you pay a price somewhere else, usually needing constant compaction on live, query handling nodes. Rbranson elaborated really nicely on this stuff:

http://news.ycombinator.com/item?id=2688848

Other DBs also win in some ways by managing their own row cache. Letting the OS do it against disk extents means holes in RAM data. If you handle it yourself, you focus on caching actual rows and not deleted data. Cassandra does this.

Once MongoDB addresses compaction, locks, and yielding, it will start to be very competitive.

Keep in mind that I think there's something to be said for having native replica sets and sharding. I can say that it definitely does work in 1.8, we use it all over the place. With the recent journaling improvements, durability is also better, but even without it our experience has been more than adequate. Secondary nodes keep up with our primaries just fine. In the rare case where a primary falls over, we lose maybe only a handful of operations, which for our workload is acceptable.

Overall, we're happy with our choice of MongoDB, it's already doing what we need. With the improvements coming down the road, I think it'll be a major force in a year from now, so I'm also excited for what the future holds.

This post is a great 101, but it doesn't provide much info on what kind of setup you run at Mixpanel. I see you ditched Cassandra, though, and you mention Riak. Are you using Riak, rolling your own sharding layer, or what?