HN user

kvh

920 karma

Founder at Patterns (https://www.patterns.app)

YC Badge: 0x400fe630c9f32773504c102910918d0e6c47f3fa

Posts16
Comments37
View on HN

Author here, I’ve updated the post. The first draft of this app and blog post took me two hours, but I kept coming back with new ideas and tweaks throughout the week. By the end, I’d certainly spent more than two hours (more like 8?), so you’re right, I just failed to update the post. The main point stands — it’s surprisingly good for the amount of effort put in (although unclear how much more juice you could get out of gpt with more effort. Clear diminishing returns)

Yes, great point, we share that concern. All of our components (patterns/openai-completion@v4) are open-source and can be downloaded and "dehydrated" into your Patterns app. They all use the same public API available to all apps.

We're working towards a fully open-source execution engine for Patterns -- we want people to invest with full confidence in a long-term ecosystem. For us, sequencing meant dialing in the end-to-end UX and then taking those learnings to build the best framework and ecosystem with a strong foundation. Stay tuned!

Thank you for the kind words and congrats on the great work on Orchest!

The marketplace is an open ecosystem, yes! Anyone can build their own components and apps and submit them. More details here https://www.patterns.app/docs/marketplace-faq/, and guide for building your own: https://www.patterns.app/docs/dev/building-components. It's early days but our goal is coverage of all data sources and sinks, the ontology layer of common transformations and ETL logic, and AI / ML models.

Those are great tools, but built for a different era. We've built Patterns with the goal of fostering an open ecosystem of components and solutions that interface with modern cloud infrastructure and the rest of the modern data stack, so folks can build on top of other's work. As more and more data lives in the cloud, in standard saas, more and more businesses are solving the same data problems over and over. We hope to fix that!

I like Matt, but that's a misleading comparison. You can't compare banks with non-banks, since, again, banks have the special regulated right to _create money out of thin air_ and put it on their balance sheet (or if you prefer the fractional reserve metaphor, they can double count your deposit -- lending it out while pretending they still have it for you).

Here's a way to reality check the difference: if 90% of Tether holders redeem their deposits tomorrow, 100% of them will get their money back and Tether Ltd will remain 100% solvent and liquid. If 90% of JPMC depositors redeem tomorrow only 10% of them will get their money back and JPMC will be insolvent.

Where would you rather have your money?

The article isn't saying what people think it's saying, but tether fud makes good clickbait I guess. Tether has indeed misrepresented its balance sheet at times, but the reality is it's a highly over-capitalized bank -- whereas most banks have liquidity ratios of ~10% (less than that pre-2008) no one is questioning tether is >50%.

A common misconception is that banks use "fractional reserve" lending, in reality private banks create money out of thin air when making loans, constrained only by regulated capitalization requirements (and the obligation to take the write-off on their own balance sheet should the loan default) [1].

Another common misconception is that unregulated banks lead to financial instability and panic. The theoretical and historical evidence for this is pretty weak [2] -- people are much more vigilant with their money when banks are unregulated, and much more aware of the inherent risks of financial systems.

(If all of our regulations worked so well, why are our financial crises worse than ever? cf 2008)

[1] https://www.bankofengland.co.uk/knowledgebank/how-is-money-c... [2] https://www.jstor.org/stable/1814673

SEEKING FREELANCER | SF | Remote or local

Looking for front-end dev (angular / typescript) to accelerate our app development. We are silicon valley veterans (Google, Square) building a data pipeline platform based on functional reactive components. Initial medium-term project to start, open to remote or local in SF.

Email me at kenvanharen@gmail.com

Hi HN, I've been doing data science for 10 years in silicon valley, snapflow is my attempt to bring the best practices of software engineering to the world of data. Concepts like modularity and reusability, pure functions, testability, gradual typing, and immutability.

The goal is a framework that makes building industrial grade data pipelines fun and fast!

What gets me most excited is the ability to share and re-use data fetchers, transformations, analysis, and models -- the foundation for a collaborative data ecosystem.

Thanks for checking it out and feel free to reach out with thoughts! kenvanharen@gmail.com

Impressive the abstraction NNs can achieve from just character prediction. Do the other systems they compare to also use 81M Amazon reviews for training? Seems disingenuous to claim "state-of-the-art" and "less data" if they haven't.

What a tragic waste of data and time. Not one mention of confidence intervals (are _any_ of these differences statistically significant??), selection bias (who was more likely to submit photos, and why did they choose a specific photo??), or sampling errors (who rated the attributes, and how consistent were they?). The OK Cupid blog posts are a great source for similar (but statistically sound) studies.

You're joking right? The fact that it even mentions the terms randomized and observational puts it in the 99th percentile of medical science reporting.

And why can't you combine results from observational studies and controlled studies? Surely they both provide evidence (albiet very weak evidence in the former's case) of the effect.

That specific point stuck out as incorrect to me as well. And given that it's one of the only technical points in the article, doesn't really inspire confidence in the author. That said, however, I do tend to agree with the overall hypothesis that the value add of "big data" is very marginal for most established industries, for exactly the reason he mentions: they've all been doing this for 20 years now. It's just become cheaper and easier recently. There are certain domains where modern data techniques can have a transformative effect, he mentions one (law). I don't think the data analytics solutions and companies that are being talked about are capable of generating these kind of transformative innovations though; so, I think the author is absolutely correct to say that the big data hype is overblown. There _will_ be big opportunities for big data, but they will come from disruptive startups finding new ways of doing old jobs. Not from crunching a few more GB of your data for a few less bucks.

This is something I've been thinking about seriously, building an "academic-level" search engine. I have the IR/NLP background. If anyone is interested discussing/collaborating, ping me!

You could just map 5 to something farther away, like 6. In fact, this is how most ordinal inference techniques work anyways: by taking an interval method and learning cutoffs for your ordered categories. Learning more parameters comes with a big cost though, which is why in practice the cutoffs are often fixed from the get-go. Obviously, ordinal methods have been tried in the literature. There is a reason they are not used in practice though, and that's because the trade-off (being harder to learn vs modelling the data more accurately) is not favorable.

>Theoretically, you should be able to get your study reviewed not just by a single expert, but by every expert on that subject regardless of geographical location or journal affiliation.

This is definitely a direction I would like science.io to go in in the future

I agree on all your points. This is my attempt at getting a foot in the door--focusing on computer scientists, who have proven more open to these kind of changes; keeping things anonymous but attributable; making the site useful (hopefully) to practitioners and folks outside of academia, etc. We'll see if it is enough.

I'd love to hear your ideas, I'll definitely follow up with an email. Thanks.

I disagree. This is another example of underestimating how powerful the human brain is: it is an exquisitely designed piece of dedicated hardware with more power than even our super-compute-clusters. Hardware that's dedicated, to a large extent, to language processing. I think current statistical approaches, ramped up on future computing resources, will continue to chip away at 'word error rate'.

The P in P(X) and P(Y) IS actually the same P. It represents the probability of the underlying sample space. X and Y are random variables mapping from that sample space to the real line. P(X=x) is shorthand for P(X^-1(x)).

For instance, the same authors, in a previous paper analyzing nigerian election results, state the following "lab experiments indicate that individuals tend to favor small numbers, even when subjects have incentives to properly randomize. Second, individuals underestimate the likelihood of digit repetition in sequences of random integers, so we should observe relatively fewer instances of repeated numbers in manipulated vote tallies." there is no mention of either of these statistics in the Iran analysis.

I think his point is that if a hundred such political scientists used a hundred different methods of evaluating election results, one of them would be guaranteed to find evidence of fraud. the other 99 would not report there results, not having found anything exciting. im not convinced there ARE that many tests you could perform, but the general point is interesting...