HN user

theboat

111 karma
Posts2
Comments30
View on HN
GPT-4o 2 years ago

I love how this comment proves the need for audio2audio. I initially read it as sarcastic, but now I can't tell if it's actually sincere.

I didn't see this in the blog post, but did you train this from scratch or finetune an existing base model?

If from scratch, quite impressive that the model is capable of understanding natural language prompts (English presumably) from such a small, targeted training set.

How does datasette work with unstructured data?

I work with large text datasets, and I typically have to go through hundreds of samples to evaluate a dataset's quality and determine if any cleaning or processing needs to be done.

A tool that lets me sample and explore a dataset living in cloud storage, and then share it with others, would be incredibly valuable, but I haven't seen any tools that support long-form non-tabular text data well.

Thank you for open sourcing this. More competition in the budding metrics ecosystem is good for end users.

It seems like you think MetricFlow should be the data mart layer and not just the metrics layer. If that's true...why? Why would I join my fact and dimension tables in metricflow instead of in dbt? One of the value adds of dbt is that it centralizes business logic in a single place. Joins are business logic. The industry seems to be moving towards creating very wide data mart tables in dbt and surfacing them to the semantic layer 1:1, or building the metrics layer on top of them.

And I wonder if a person free of cognitive distortions would even be referred to as human, as the quote goes:

There's a difference between emotional intuition and emotional reasoning (the cognitive distortion in OP's example).

Emotions are extremely valuable for decision-making (e.g. this house ticks all my boxes but do i love it?) and making judgements (e.g. this situation does not feel right to me).

Emotional reasoning is when people distort reality in favor of their (often self-destructive) emotional impulses, discarding physical evidence in favor of their emotions.

Does this imply airbyte only supports connectors that airbyte validates and integrates into the platform? Can I use an airbyte connector that lives in a repo on my private github?

This solves the problem of getting high quality connectors built, but how do you plan to maintain them? What if the original contributor falls off the face of the earth?

Sometimes it feels like we're in the midst of a loneliness epidemic, created or exacerbated by technology like the internet. But the data doesn't actually line up with that belief: https://ourworldindata.org/loneliness-epidemic

It looks like we are generally as lonely as we were throughout the 20th century. Perhaps that's a sign the internet hasn't yet lived up to its promise as the great unifier. The fact that the greatest and most accessible communication technology in history hasn't put a dent in loneliness shows that we still have a lot of work to do in making it serve that end.

I once thought community-based mesh networking could disrupt last-mile internet providers, beyond its obvious value in remote areas and disaster zones. If we could get decent bandwidth with a freemium/pay-as-you-go model, it would be way more consumer friendly and affordable than paying a big telco to string a wire into your house.

However with the advent of 5G and now satellite internet, it seems like high-speed wireless internet will be ubiquitous in relatively short order without the need for mesh networks. So that dream is probably dead.

Making use of airflow's plugin architecture by writing custom hooks and operators is essential for well-maintained, well-developed data pipelines (that use airflow). A little upfront investment in writing components (and the most painful part, writing tests) will go a long way to helping data engineers sleep at night.

That said, I make a point of using ETL-as-a-service whenever it's available, because there's no use solving a problem someone else has solved already.

I would make use of ETL-as-a-service vendors as much as possible (Stitch, fivetran, etc.) before going to airflow/prefect for the more custom stuff.

Airflow 2.0 will have some pretty nice features for ML development as well.

I'm a Canadian with a BA in economics, but I now work as a data engineer at a startup. I see postings for remote data engineering jobs in the US which say that I must be eligible to work in the US and/or have a social security number.

Is it possible for me to work for one of these companies using a TN visa? Does my degree have to be "engineering" if the role is "data engineer"?

It's important to note that remote work in March/April/May (and perhaps even today) was uniquely coupled with remote life, since most knowledge workers stayed home due to social distancing guidelines. It should not be surprising that the boundaries between work and life blurred when life itself came to a standstill.

I'd be surprised to see if the same tendency to blur life and work, such as by working on evenings and weekends, persists once life returns to its previous form.

If the web migrates to biometric sensors for authentication, I hope this won't suffer from vendor lock-in. When every new device ships with facial recognition and/or a fingerprint reader, it will be nice to login using my face/fingerprint irrespective of the device I'm on.

There have been so many opportunities for a grassroots pro-privacy movement to develop, and yet there isn't one. Devastating hacks (Target, Yahoo), election interference (CambridgeAnalytica), and yet nothing.

Acting as if people are unaware of data collection is disingenuous. If you told the average facebook user how much facebook and its third-party partners knew about them, I doubt many of them would stop using the platform.

While I sympathize with your viewpoint, you're imposing your personal values on consumers who demonstrate their willingness to exchange personal data for free services every day. You and many HN users may balk at this, but most people are ok with trading privacy for real-time traffic predictions. Apple shouldn't receive an unfair market advantage because they embody the values you hold dear.

Sidenote: I disagree that Apple Maps' success puts pressure on Google to up their privacy game. On the contrary, Google Maps comparative advantage is their data trove, as there are many more users of Google Maps than Apple Maps, so they seem more likely to lean on that to succeed.

I wouldn't look to the market to improve privacy, since as I said above, the market clearly doesn't care about privacy much at all. Without a seismic shift in public attitudes towards privacy, it's up to the government or the companies themselves to adapt.

As much as I agree with your general sentiment, Apple Maps, like many of Apple's mobile apps, gets a boost from Apple's anti-competitive practices. It's utterly ridiculous that we can remove default apps from iOS, including Maps and Safari, but we can't set new default apps to replace them.

If we ever get serious about increasing competition in the tech sector, an easy place to start is letting users set default browsers, maps, and email clients on their devices.

When everyone accepts libra, there is no need to exchange libra for hard currencies. Facebook, as the owner and operator of the world's largest social network, believes it can create a closed ecosystem where libra is used for many types of transactions.