HN user

_hl_

511 karma

hn@henrikl.com

Posts5
Comments177
View on HN

Hey, I’m acutely in the market (considering moving away from Google)

2 Qs:

1. How does OpenCage correctness/completeness compare to Google Maps API, especially in rural and industrial regions where you have addresses like “AcmeCo Industries, 234-XY Unit C, Jebel Ali Free Zone, Dubai”? I’d like to confidently query the most precise location that still matches/contains my query.

2. Do you support querying by business names? Google’s geocoding doesn’t return the business name in the result (that’s a separate API), but it does use business names to resolve queries.

I encourage anyone to read the (surprisingly plain-english) first few pages of the decision, but here is the gist of it:

In January 2024, the court issued a post-trial opinion finding that the award was subject to review under the entire fairness standard [...] the defendants bore the burden of proving entire fairness, they failed to meet their burden, and the plaintiff is entitled to rescission. [...] The defendants responded by putting the rescinded compensation plan [...] to a stockholder vote for the stated purpose of 'ratifying' it. [...] The defendants then moved to 'revise' the post-trial opinion based on the stockholder vote, asking the court to flip its decision.

The motion to revise is denied. [...] The large and talented group of defense firms got creative with the ratification argument, but their unprecedented theories go against multiple strains of settled law. [...] First, the defendants have no procedural ground for flipping the outcome of an adverse post-trial decision based on evidence they created after trial. [...] Second, common-law ratification [...] cannot be raised for the first time after the post-trial opinion. [...] Third, [...] a stockholder vote standing alone cannot ratify a conflicted-controller transaction. Fourth, [...] material misstatements in the proxy statement [defeat the ratification]. Each of these defects standing alone defeats the motion to revise.

The fee petition is granted in part. The plaintiff’s attorneys asked for $5.6 billion in freely tradeable Tesla shares. [...] That was a bold ask. [...] Delaware courts award fees based on a percentage of the value of the benefit achieved [...] yet [...] a fee award 'can be so large that typical yardsticks [...] must yield to the greater policy concern of preventing windfalls to counsel.' [...] $5.6 billion is a windfall no matter the methodology used. [...] To reach a reasonable number, this decision [...] uses the $2.3 billion grant date fair value to value the benefit achieved. [...] Applying a conservative 15% to that figure results in a fee award of $345 million—an appropriate sum to reward a total victory.

For anyone else confused by what actually happened, here is a summary compiled from various sources around the conviction and the related 1MDB scandal:

---

Jho Low, a Malaysian financier, masterminded one of the largest embezzlement scandals in history through 1Malaysia Development Berhad (1MDB), a sovereign wealth fund intended to spur economic development. Over $4.5 billion was siphoned from the fund to finance a lavish lifestyle, high-profile investments, and extensive political influence campaigns. Fleeing justice in Malaysia, Low focused on cementing his power in the U.S., including efforts to influence the political landscape and suppress investigations into his crimes.

Pras Michel, a founding member of the hip-hop group Fugees, became entangled in Low's schemes, leading to his conviction on 10 criminal counts. Michel first met Low in 2006, and by 2012, he was a key player in Low’s efforts to use his ill-gotten wealth to influence U.S. politics. Low funneled $20 million to Michel to gain access to then-President Barack Obama’s re-election campaign. Knowing direct contributions from foreign nationals were illegal, Michel orchestrated a scheme using straw donors and political committees to route Low’s money into the campaign. Michel also used funds to buy seats at fundraising events and pressured wealthy acquaintances to participate.

By 2017, Michel’s involvement deepened as he acted on behalf of both Low and the Chinese government without registering as a foreign agent. In exchange for millions, Michel attempted to influence the Trump administration to drop the U.S. investigation into Low and to extradite Chinese dissident Miles Guo, a target of Beijing. These actions violated federal law, which requires registration for such foreign lobbying efforts.

Michel was also convicted of laundering millions of dollars tied to the 1MDB embezzlement and attempting to obstruct justice by pressuring straw donors to support his version of events during the investigation. The trial revealed Michel’s use of burner phones to contact witnesses, an act he later admitted was misguided. His defense argued that Michel was unaware of the legal boundaries and acted on bad advice from his attorney, including the use of artificial intelligence to craft his closing argument—a controversial decision.

The prosecution presented Michel as a knowing participant in a broader conspiracy to influence U.S. politics and aid foreign interests. Testimony from high-profile witnesses, including actor Leonardo DiCaprio and former Attorney General Jeff Sessions, underscored the scale of the scheme. Michel was ultimately convicted of conspiracy, campaign finance violations, acting as an unregistered foreign agent, money laundering, and witness tampering.

The frustrating thing with SOC2, or pretty much most compliance requirements, is that they are less about what’s “technically true”, and more about minimizing raised eyebrows.

It does make some sense though. People are not perfect, especially in large organizations, so there is value in just following the masses rather than doing everything your own way.

Read the notes in the link you posted. I don’t think it says what you think it says.

In May 2020, the definition of M1 (monetary supply in “cash”) was changed to include savings deposits. They changed this not due to some conspiracy, but because savings accounts were deregulated to remove withdrawal limits, effectively rendering them cash-equivalent, and thus necessary to include in M1 metrics.

I.e. the 80% spike has nothing to do with money being printed.

I mean a normal passenger on a normal plane making a normal trip to an office building and finding a hidden location where to tape a small box with an arduino in it. Maybe even on the outside so you can use solar power? Though it only needs to last long enough to compromise a machine inside the network.

This would be nothing new, I remember ages ago in the days of WEP that you could buy a small box that would collect enough handshakes to let you crack the WEP password.

What’s wrong with the tried-and-tested technique of flying a guy or girl over there to drop a small gadget in WiFi proximity?

You’d need to go a level below the API that most embedding services expose.

A transformer-based embedding model doesn’t just give you a vector for the entire input string, it gives you vectors for each token. These are then “pooled” together (eg averaged, or max-pooled, or other strategies) to reduce these many vectors down into a single vector.

Late chunking means changing this reduction to yield many vectors instead of just one.

In a perfect market, the market maker who sells you that option offsets it with correlated assets in the other direction, eg by buying or selling stock that is sensitive to the election.

Large trading firms exist on finding and exploiting small arbitrages between various correlated assets. If you assume a perfect market with infinitely many participants and infinite liquidity, then this “works” - there is no distortion at scale.

This is awesome, but I'm not sure what the long-term use case for the intersection of low-latency integration and non-production-stable is? I'm saying this as someone with way more experience than I'd like to in using reverse-engineered APIs as part of production products... You inevitably run into breakages, sometimes even actively hostile platforms, which will degrade user experience as users wait for your 1day window to fix their product again.

Though I suppose if you can auto-fix and retry issues within ~1minute or so it could work?

I see - I suppose that’s a fair viewpoint to have!

I’m not much of a python programmer but my experience with the language would make me tend to agree actually. There are bigger fish to fry and so the effort to go after this relatively tiny sardine is perhaps not worth it.

1. Slack has network effects: we connect with customers on Slack (or M$ Teams for enterprise…)

2. I don’t want to innovate on “back office”. Slack works, is sufficiently affordable, and costs no social credit with employees.

3. I know I won’t run into problems in the future. This kinda ties into (2), I don’t want to innovate on back office, but to make it concrete: Deel, Rippling and the other M$ AD clones all integrate with Slack to set up permissions and SSO with zero effort.

4. Slack has lots of sensitive info about us and pur customers, which makes it SOC-2 relevant. I want to use “industry standard” tech for anything compliance related. Though I don’t recall that this would have ever been a problem on security questionnaires, and not many years ago Slack was the “young kid on the block” themselves & managed, so idk if this is actually a valid point.

If you give me a feature parity clone of slack for half the price, I’d certainly switch, but anything less than feature parity and I probably wouldn’t. I don’t need to or want to take risks on internal tooling.

My understanding was that the extra parameters required for the second attention mechanism are included in those 6.8B parameters (i.e. those are the total parameters of the model, not some made-up metric of would-be parameter count in a standard transformer). This makes the result doubly impressive!

Here's the bit from the paper:

We set the number of heads h = dmodel/2d, where d is equal to the head dimension of Transformer. So we can align the parameter counts and computational complexity.

In other words, they make up for it by having only half as many attention heads per layer.

Some of the "prior art" here is ladder networks and to some handwavy extent residual nets, both of which can be interpreted as training the model on reducing the error to its previous predictions as opposed to predicting the final result directly. I think some intuition for why it works has to do with changing the gradient descent landscape to be a bit friendlier towards learning in small baby steps, as you are now explicitly designing the network around the idea that it will start off making lots of errors in its predictions and then get better over time.

This certainly contributed to people preferring card over cash, making merchants loose ~3% per transaction.

That ship has long sailed, but it does male you wonder: if everything was priced at increments of, say, quarters, would enough people still use cash to offset the lost sales from the allegedly less appealing pricing?

I see what you're saying, but I don't think it applies in this case. Correct use of jargon helps domain experts communicate with higher precision, and papers tend to be written by domain experts for consumption by other domain experts.

Of course there are some (possibly many!) papers where jargon is abused to make something sound smarter. Sometimes this can also happen unintentionally.

In this case, "compute-optimal X" is standard terminology used in large-scale ML model design for finding the most optimal tradeoff with regards to compute when trying to achieve X.

Here, the paper is about finding the optimal model size tradeoff when training on LLM-generated synthetic data. Imagine you have a class of LLMs, from small to infinitely large. The larger the LLM, the higher the quality of your synthetic data, but you will also spend more compute to generate this data ("sampling" the data). Smaller LLMs can generate more data with the same compute budget, but at worse quality.

The paper does some experiments to find that in their case, you don't always want the largest possible LLM for synthetic data (as previously thought by many practitioners), instead you can get further by making more calls to a smaller but worse LLM.

Does that image look like a render to anyone else? Not doubting it’s real, just confused- esp. if you compare it to the photos of raptor 1 and 2 linked in another comment. Maybe it’s the background scene that makes it look too perfect?

Re. postgres, this is actually something I have always struggled with, so would love to learn how others do it.

I’ve only ever worked in very small teams, where we didn’t really have the resources to maintain nice developer experiences and testing infrastructure. Even just maintaining representative testing data to seed a test DB as schemas (rapidly) evolve has been hard.

So how do you

- operate this? Do you spin up a new postgres DB for each unit test?

- maintain this, eg have good, representative testing data lying around?

I don’t think “DB OS” is a standard term, but what I believe the commenter means is that postgres has evolved to become an incredibly mature, well built set of DB building blocks that are exposed via extension APIs etc to implement whatever kind of database you want.

Sure, there’s the core row-major RDBMS you get out of the box, but could easily turn that into a distributed column-oriented analytics data warehouse if you wanted to.

It is your responsibility to look after your team as a whole. Sometimes that means taking decisions that negatively affect a single individual.

Your team members want to work with other great people who can help them deliver success - after all, the feeling of success is one of the most rewarding things a job can give you.

I’ll also add that letting someone go (constructively) can be a net positive for them, if you can put them on a path for greater success in life elsewhere. Struggling to deliver for months or even years is also not a good use of their time.

It might feel harsh, and it’s definitely a very difficult conversation to have, but I think if you remember that your responsibility is towards the team as a whole it can help you emotionally accept and navigate the situation.

I strongly agree with the premise that an orchestrator-centric approach is preferable to event spaghetti for a lot of business process use cases.

Another (maybe more mature) alternative to Infinitic is Camunda, who have been pushing this rhetoric for over a decade.

My problem with both is that neither feel very modern, in the sense that they are tied to (in Infinitic's case) or strongly favor (in Camunda's case) Java-centric development, don't have a great developer experience story, and don't feel very cohesive with the language.

What I'd love to see is someone tackling this for smaller orgs where the above tradeoff in developer experience isn't worth the gains, i.e. orgs with <100 engineers. Something where you have a single source of truth for your processes and schemas, with version control, that directly yields strongly typed integrations & hooks into any language, with a great local developer experience and deployment story.

Camunda requires too many magic strings and schemas to be kept consistent across services for that, and Infinitic forces me to use Java or Kotlin which I don't want.

A "simple" solution might be to have declarative process and schema definitions in some DSL for version control, which auto-generates schemas, types etc in whatever language, giving me both strong types, intuitive local development, and clear deployment story through my existing CI/CD.

There is still an unfathomably huge landscape of problems where C++ is the more mature, “right” choice over rust for new projects. Just from my (extremely limited and narrow) personal experience:

- scientific/numerical computing (super mature)

- domain-specific high performance computing (C++ is for better or worse extremely flexible)

- weird CUDA use cases (feels very native in C++)

- game dev as you mention

- probably much much more

C++ to me feels more like a meta-language where you go off to build your own kind of high performance DSL. Rust is much more rigid, which is probably a good thing for many “standard” use cases. But I’m sure there will always be a very very long tail of these niche projects where the flexibility comes in handy.

People here seem to treat this like advertising, because it kinda sounds familiar to advertising. I’m as critical of ClosedAI as the next guy, but let’s think that idea through: OpenAI are the ones paying the content provider for exposure, not the other way around. In return they get training data.

The only reason for OpenAI to do this is if it makes their models better in some way so that they can monetize that performance lift. So I think incentives here are still aligned for OpenAI to not just shill whatever content but actually use it to improve their product.