HN user

vkb

1,536 karma

Vectors. Distributed systems. Data.

vickiboykis.com @vickiboykis.com on bluesky Veekaybee on GitHub

Posts71
Comments71
View on HN
vicki.substack.com 6y ago

All numbers are made up, some are useful

vkb
6pts0
www.vox.com 7y ago

Many books aren’t fact-checked, and we’re realizing they’re full of errors

vkb
14pts3
www.curbed.com 7y ago

Can money create a neighborhood?

vkb
1pts0
metatalk.metafilter.com 8y ago

Metafilter financial update and future directions

vkb
1pts0
technicaldiscovery.blogspot.com 8y ago

Reflections on Anaconda as I start a new chapter with Quansight

vkb
2pts0
newsroom.fb.com 8y ago

Facebook Newsfeed: Bringing People Closer Together

vkb
3pts0
tanmulabs.com 8y ago

Introducing the Zarf beta

vkb
1pts0
www.chronicle.com 9y ago

Lyft and learn

vkb
1pts0
ironedcurtains.com 9y ago

Eye-witness stories from Chernobyl

vkb
154pts121
news.ycombinator.com 9y ago

Ask HN: What are we doing about Facebook, Google, and the closed internet?

vkb
609pts424
www.reddit.com 9y ago

How I Lost $7 Selling McDonalds Sauces Online

vkb
2pts0
education.penelopetrunk.com 9y ago

What is computer literacy in today’s world?

vkb
2pts0
lists.apache.org 9y ago

“Big data” problem(2016)

vkb
2pts0
veekaybee.github.io 9y ago

What should you think about when using Facebook?

vkb
363pts234
medium.com 9y ago

Dealing with countries – in data

vkb
3pts0
legalaffairs.org 9y ago

From Russia with LØpht (2002)

vkb
1pts0
www.racked.com 9y ago

The Case for the American Mall

vkb
1pts0
www.reddit.com 9y ago

How should I store data that I want to be able to use in 50 years?

vkb
6pts0
www.themillions.com 10y ago

There Is No Handbook for Being a Writer

vkb
161pts80
www.vulture.com 10y ago

Lit-Blog Pioneer Jessa Crispin Closes Bookslut

vkb
3pts0
veekaybee.github.io 10y ago

Content is dead

vkb
75pts109
well.blogs.nytimes.com 10y ago

Don’t Post About Me on Social Media, Children Say

vkb
2pts1
arstechnica.com 10y ago

IBM’s 1937 corporate songbook (2014)

vkb
2pts0
www.businessinsider.com 10y ago

Chick-fil-A debuts new valet service for parents of young children

vkb
1pts0
www.nytimes.com 10y ago

Parents Ready for Some Love from Silicon Valley Companies

vkb
14pts14
www.newrepublic.com 10y ago

Most writers have no choice but to bend the rules

vkb
2pts0
www.theatlantic.com 10y ago

The dog-walking economy

vkb
15pts3
www.theatlantic.com 10y ago

The dog-walking economy

vkb
1pts0
www.philly.com 10y ago

Philadelphia network gets upgrade but still may not bear pope traffic

vkb
1pts0
www.theguardian.com 10y ago

Inside the Red Web: Russia's back door onto the internet

vkb
4pts0

Author here. That's definitely my bad and not an intended user experience. The text was initially meant as a transcript accompanying presentation slides. I've compressed the images so it should be at least slightly better now.

The good news is that this already exists in the Redis search module [1], which allows you to do similarity search against indexed embeddings, among other features, and offers comparable performance against other ANN libraries [2], depending on your performance criteria.

I've been using it for a side project to do semantic search on books[3] and have been really happy with its performance. (Not affiliated with any of this, was mostly interested in exploring existing well-performing, fairly standard tools with low latency)

[1] https://redis.io/docs/interact/search-and-query/search/vecto... [2] https://ann-benchmarks.com/#redisearch [3] https://viberary.pizza/how

Author here. The primary difference is that Confluence is more like a reference manual. A P2 is more like a conversation and a record of decisions made, along with all the context for those decisions, since each post has a threaded comment section where people can reply to each other. It's not better or worse, but has a slightly different use-case in that it's essentially like email threads or RFCs that anyone in the company can read.

Editor here. Thank you for sharing your memories. And, for the note.

Some details about the origin: the vignettes were originally English-language comments in a private Facebook group.

It was a huge team effort to get them all collated, organized, and to check and re-check attribution and consent to share. We are so happy that we are able to share these collective memories and eyewitness testimonies, particularly outside of Facebook's walled garden.

Editor here. These were originally written in English, so we have plans to translate to Russian, which might take a while since we only have one fully-qualified translator. Help is welcome! Shoot me an email (included in my profile).

This is a rather interesting stance to take for a publication whose article quality has degraded and ad size has increased (thirty-eight trackers, including two from Facebook, detected by my adblocker when I tried to access the article) over the past five years to the point where I refuse to read them.

If you take a look at the homepage through the Wayback archive as it used to appear in 2005[1], 2011[2], and today, you'll see how content disappears and click-baity headlines rise over time.

The Atlantic is very much a part of the problem of "the race towards the bottom" the author describes, and instead of having a discussion about how to fix it and maybe trying different revenue models, it continues to un-ironically have share and tweet buttons at the top of this article.

[1] https://web-beta.archive.org/web/20050210070148/theatlantic.... [2]https://web-beta.archive.org/web/20110731234233/http://www.t...

I don't know that that's necessarily true. The most recent StackOverflow survey[1] shows a difference of 8%, which is not an overwhelming majority. Granted, that's not an unbiased sample size, but I think the OP above is correct...more data scientists use Python than Java.

So anyone wanting to use this library would have to think about tradeoffs: Are the efficiencies lost in data scientists learning to use Java for modeling worth the efficiencies gained in putting a model in production? For some, the answer may be yes, for some no.

[1]https://stackoverflow.com/insights/survey/2017#technology-pr...

Perhaps it's just my selective biases, but I feel like I am reading about more notable people in the tech industry speaking out in favor of privacy over the past couple months than even when the Snowden revelations happened in 2013.

Could this possibly be a tipping point for adtech as a revenue model? (Although I've been reading about it for almost a year now [1], [2], [3]) I'd like to hope so, and am also curious: how invasive can companies be before consumers start to push back?

[1] http://digiday.com/marketing/copyranter-ad-tech-bubble-explo... [2] https://www.reddit.com/r/investing/comments/41zyoo/why_i_bel... [3] https://kalkis-research.com/google-end-of-the-online-adverti...

Data scientist here. It is 100% possible to do things with kids, but you really have to be motivated to do it, AND it really helps if you have a support system of other people to help take care of your child. I wrote about the dangerous deception we have in American culture, and particularly tech culture, of people who "have it all," but in reality have a bunch of help in the background here[1] and here[2].

If you work full-time and you want to go above and beyond, you're essentially working three days: work, before school + after school, and then your third day is learning or development.

Whatever that means for you in terms of reshuffling energy and other commitments will vary on your personality, energy level, etc, etc.

When you have a small child, it is extremely hard to multitask. So I wait until she is asleep. All after work time and weekends are for her.

Here is the way my schedule works: I pick her up from daycare, do dinner, playing, and then she goes to bed. I then take half an hour break, and delve into whatever I have going on, for about three hours.

I'm currently taking a Java class, writing technical blogs, and working out some Python. So I'll usually do an hour of reading/Java homework, then start a blog post, then finish off with whatever else I was working on.

Over the past three weeks, I developed this talk on big data[3]. That was probably the hardest because I needed a lot of time to write the code, test the code and concentrate, and all of my energy was just sapped.

All of this is to say that you can do it. For me personally it takes a lot of reshuffling and work and giving up things, but that's how kids work.

[1]http://blog.vickiboykis.com/2015/09/we-are-not-getting-the-f... [2]http://blog.vickiboykis.com/2012/07/sheryl-anne-marie-and-ma... [3]https://veekaybee.github.io/data-lake-talk/#/

R for Excel Users 9 years ago

There are definitely a lot of people who fall into those categories, but the articles/links give the impression that everyone is senior. Which is why it's great when articles like this come around.

R for Excel Users 9 years ago

This is a fantastic article for intermediate beginners. On HN, everyone is a senior data scientist working with Spark and Keras and Tensorflow and deep learning.

In the real world, there is a huge chasm of difference between people just learning Excel and developers, not many people even understand why you would switch away from the former when it's so convenient, which is why the difficulty v.s. complexity chart is so great, and may actually speak to people in an approachable way.

There are a lot of tutorials for how to do hard things and how to do easy things, but not a lot for how to think of the hard things in terms of the easy things, and this falls in that category. Another good book on this topic is Data Smart by John Foreman, where he goes over basic data science skills in Excel.

I am disappointed at the type of criticism of this post on Hacker News. This write-up is the kind of high-quality, original, analytical writing we so desperately need these days, in an online world that is completely saturated with clickbait that adds nothing to our understanding of the world.

Is the piece lacking in some semantics? Perhaps. But I was struck by the comments about market share and Android strategy, rather than directly discussing the article and the points it makes about the two maps at hand.

If your wife is already interested in writing about this, and already doing it, then there is absolutely value: to her. She gets to process what she's working through and pour it out on paper. That's already great. It will absolutely also look great on her portfolio. It might also be good tech exposure for her to start working with either Wordpress, or extra bonus points, Jekyll, and version control.

Speaking as someone who mentors other people in making their way through data to programming, yes, please, please have her write about this!

There are so many people looking to get into programming, but literally have no idea where to start because the whole universe is overwhelming to them and full of people who seem to be programming forever. HN is probably not the audience for this blog, but hundreds of thousands of data analysts and people who use Excel on a day-to-day basis are.

Finally, a nitpick, but I'd take issue with "These aren't enough to actually work in our industry" - if she's working with technical skills and wants to learn more, she's already in the industry and then some.

Good luck to her!

The more I read, the more I realize that fiction has a way of teaching the same lessons as nonfiction, only in a better, humane way that is not like someone lecturing at you from a podium, but more like a friend sitting down for a heart-to-heart conversation with you.

In this vein, I've started catching up on all the classics I haven't read yet, and there are a lot of them. At first, I was hesitant to dive into classics, because I thought they were hard to read and highbrow. This has mostly not been the case (with the exception of Wuthering Heights, which I found impossible to get through).

To wit, the best books I've read this year include:

A Tale of Two Cities - Impossible to get through the first several chapters, but after that, you're off to the races between France and England in the 18th century. Dickens used historical books as reference, but recreated the mood entirely from imagination. A better primer than anything in the news about Anglo-French relations, learning trust, and sacrifice.

No Country for Old Men - Want to understand the current fear of the Mexican border that has so bolstered Trump's popularity, as well as the ramifications of having a country full of veterans who don't have any medical support or care? Read this book.

Farenheit 451, A Clockwork Orange, Handmaid's Tale - Much better and scarier than any news clipping today about the possible future path for government interference in thought and action. Takes things to their logical conclusions.

And one non-fiction book that I did enjoy, Bringing Up Bebe, which is about the contrasts between child-rearing in France and America, but on a much larger scale, about the different things that seem culturally obvious to us, but are completely different to other cultures. This book made me really re-examine American food culture in a way Michael Pollan's diatribes have not.

Content is dead 10 years ago

What's so difficult about just opening an incognito window and suffering the couple of seconds it takes for a computer halfway across the world to deliver you content while you sit at your desk?

The whole point of the post was that I'm doing something inordinately ridiculous to access a little bit of good content. Scraping and incognito browsing are both anti-patterns that are symptoms of a sick media industry. A media that is sick cannot provide us with good news and content that we can use to further our critical thinking, and we should be worried and thinking about how we can possibly solve this problem instead of trying to bypass it.

Content is dead 10 years ago

Thanks for reading. The Economist ads are a great point that I didn't address, and makes me even more worried for the high-quality news and content industry than I initially noted.

The really interesting takeaway from this article for me was that product executives make many decisions based on simple intuition, industry experience, and "I've heard it worked over there."

There are numerous mentions of failed product roll-outs(News Feed as a failure at first, Paper, etc.) But the key quote is here, when they talk about the meeting: "Cox agrees. 'My intuition, which we could prove wrong, is people just want more stuff.'

I'm sure that Facebook uses data to back up some of its assertions; after all, they did mine comments to come up with the most common reactions. But how many people's entire internet experience (i.e. consumption of news through the News Feed) is impacted just by some senior guy in a room going, "This sounds right to me," and everyone scrambling to work on what he proposes?

The stark difference in tone between the beginning of this article and the end was startling. Judging from the set-up, I thought it was going to be a profile of a company working to improve conditions in rural areas, or for immigrants.

But the inherent mismatch between the depiction of Tran's very hard early life resulting from the aftermath of war and logistical difficulties in a third-world country made the problem he's facing now, how to make food for middle-to-upper-class households in America, seem laughable and trivial.

What problems are start-ups like Munchery trying to solve? Not the conditions that created Tran and many like him, but the issue of those who don't have home chefs.

I'm not saying clawing back time from the day and isn't important and isn't valuable. It most definitely is, especially for parents working in the "second shift"[1], after work.

But depicting the idea that Munchery is some great force with a vision while ignoring the larger systemic problems in the story (poverty, access to clean water, aftermath of war) that propelled Tran to arrive in America is a failure of journalism.

Let's work on - and shed light on - the harder stuff.

[1] https://en.wikipedia.org/wiki/The_Second_Shift

I don't have any specific studies, but I have read about human cognitive biases in general (Thinking Fast and Slow, Freakonomics), have talked to pregnant women going through the interview process, and have myself interviewed as a pregnant woman over the past couple years. I've also read countless forums (Corporette posts[1], the excellent Ask A Manager[2] blog, Laurie Ruettimann [3] that talk specifically about workplace psychology in a non-clinical way.

The hiring manager likes to think they are not biased, but they always are, in the same way all humans are. Be it by names in resumes [4], or because employees experience a "cultural mismatch", we are imperfect as humans and will judge a person just as much as we judge credentials or resumes.

The best hiring managers will recognize that this is a Thing That Exists and try to actively overcome these biases. But the world is not made up of best hiring managers. It is made up of people who may or may not assume that because you are pregnant you will not do as much work, or they don't want to give you leave or deal with FMLA, or they don't want to deal with you potentially not coming back after your leave is over.

This is a constant fear for women, this fear of being judged on our childbearing potential, and comes up over and over again in both motherhood forums and professional work forums, as well as more recently in the national conversation with the likes of Lean In, Anna Marie Slaughter, and others.

I've had conversations with women who are worried that they should remove their engagement and wedding ring in case the employer sees it and assumes the woman is newly married and wants to have children right away, ruining the amount of time she might put into the career. Other women are terrified of losing their jobs after coming back from a maternity leave where men or non-pregnant coworkers handle all of the business and they suddenly feel they are not needed.

To anyone unfamiliar with how hiring and workplace psychology work, these concerns seem far-fetched. To many women working today, they are an unfortunate reality, although one that doesn't exist everywhere and is starting to shift, hopefully, with things like an increasing number of more flexible parental leave policies.

So, don't mention anything until you're far along enough in the process that it's time for both parties to evaluate whether they want to work with each other (i.e. the offer stage). Then it's 100% worth putting on the table.

[1]http://corporette.com/ [2]http://www.askamanager.org/ [3] http://laurieruettimann.com/ [4]http://freakonomics.com/2013/04/08/how-much-does-your-name-m...

How far along are you? If you're not far along enough to show, don't bring this up. Regardless of how open-minded hiring managers are, it will definitely unconsciously bias a decision. Don't make it a deciding factor.

If you have a great interview and an offer, that's the time to bring up any potential time off or accommodations that you'll need. Corporette has some really great resources around this:

+ http://corporette.com/2011/06/16/taking-a-new-job-while-preg... + http://corporette.com/2012/11/20/how-to-negotiate-future-mat...

And, congratulations!

Thank you for the background. I am super-familiar with the Soviet WWII documented experience (Vasiliy Grossman, Ehrenburg, etc. ) but I've embarrassed to say I've never heard of her.

Having worked in some form BI for almost the entirety of my career now, there is not a single, consistent form of BI dashboard that is prevalent across any company. Every solution ends up being unique, because every company has a unique data set-up, stakeholders, definition of metrics, and access needs.

I've worked with Tableau, Domo, Oracle products, you name it. What's the solution that is passed around the most? Excel sheets, because they travel easily and have all-around permissions.

I've been waiting for an out-of-the-box solution that's at least relatively easy to leverage across different organizations, but I haven't seen a painless one yet.

I'm hopeful that Quicksight, while not the be-all end-all solution, provides an example for others to follow, if it does end up being easy to set up and use.

Congrats on considering moving to the most awesome city in the world. ;)

The Philly tech scene is small but dynamic and growing every day. For a sense of the scene, read Technically Philly and subscribe to the Philly Startup Leaders listserv [2]. The major clusters are Old City Philadelphia, home to N3RD Street and other related makerspaces, Center City, University City (near UPenn/Drexel), Northern Liberties as was mentioned. There are a bunch of suburban tech clusters as well, and the city is pretty evenly split between the two, I'd say.

There are tons of meetups in any given week on anything ranging from VR to Android to data. There are a couple of big conferences here every year including Emerging Tech for the Enterprise [4]. We also just hosted FOSCON. We have an active geek community in Geekadelphia [5].

Big companies (depends on what you think of as big): DuckDuckGo, Curalate, Linode is relocating here, Comcast is the number one employer in the region and is increasingly working on more cutting-edge tech projects (although obviously also still has all the features of a large corporate entity), Urban Outfitters down in Navy Yard. Then there are plenty of pharma , financial, and eds that make up the meat of the city's white-collar industry: GlaxoSmithKline, PNC, Temple, Drexel, Children's Hospital of Philadelphia, etc.

Where to live? A loaded question depending on what you value. Some people live in the city to be close to everything. The coolest neighborhoods right now are Northern Liberties and Fairmount, probably. Some people live in the burbs for the schools. Some burbs are closer and more accessible by train than others. Depends on what you're looking for.

Any more details, feel free to shoot me an email. It's my profile. I'm happy to answer questions!

[1] http://technical.ly/philly/ [2] http://phillystartupleaders.org/ [3]https://en.wikipedia.org/wiki/N3RD_Street [4] http://phillyemergingtech.com/ [5] http://www.geekadelphia.com/

I've been thinking about this issue for a long time, both on my own, and when we diagram these models out in business school, and I haven't been able to come up with a model where both content-producers and consumers win. Their needs are directly at odds with each other.

This is probably because we are thinking about this the wrong way. The news sites today are engineered to catch your attention. But the market rewards companies who keep it. The New Yorker, The Economist, and NYT come up again and again on this site and others as examples of "the only news source I pay for" because of their quality content. Simply put, they don't make you feel stupid when you read them.

The problem is the amount of attention required to keep someone's attention is much less than the amount required to catch it, which is why BuzzFeed exists. What do people usually read on their lunch breaks? HuffPo. Until that changes, the landscape will be what it is.

I'm extremely interested in how Aeon.co will make those choices. They manage to consistently produce interesting, relevant, long articles but have no ads (yet). In an interview from 2012[1], its cofounder says, "Our business model is to spend the first year or so investing significantly in the magazine in order to build up a strong following or community of interested people – readers, writers, artists and photographers. Once we have established that reach we will start to build opportunities for generating revenue. We are not sure what forms these will take and are watching closely how other publications are doing so – from micro-payments for articles, to higher levels of service for subscribers, live events, and online fora." This quote tells me that even those in the industry have no idea and everyone is still making it up as they go along.

[1]http://www.frostmagazine.com/2012/10/brigid-hains-on-the-lau...

Is the tradeoff then visibility and platforms or obscurity and freedom? It seems to me like there's a middle ground somewhere.