HN user

jmcminis

45 karma

I’m an engineer who works on data and loves research. jeremy at sawtoothdata [.com]

Posts7
Comments33
View on HN

Are there edge cases here due to context length?

1. I have a json schema with required fields. I complete the json, but do not include the required fields.

2. I run out of token from the model before I finish the json object because I'm in the middle of some deep, nested structure.

These seem solvable, just edge cases to control for by either reserving tokens, randomly generating required tokens until completing the json, or something more sophisticated.

Here are a few things that might help

General advice:

People + AI Guidebook - A toolkit for teams building human-centered AI products. https://pair.withgoogle.com/guidebook/

LLM advice:

Tips to improve prompt and answer quality. https://github.com/openai/openai-cookbook/blob/main/techniqu...

I wrote a short overview of some of the LLM Application development tools and platforms that might be helpful: https://mcminis1.github.io/jekyll/update/2023/01/23/llm-land...

In general the CEO has the ability to award equity however they want. It's fairly common for exercise windows to be extended, awards to be increased as a part of a severance package, cashless exercise to be allowed, etc. You may need to get sign off from the board, but that's usually easily justifiable.

As stated above, after that, the communication to the rest of the team is key.

I think there is a trend towards using SQL to do the T part of ELT. For example, see the rise in popularity of dbt. Analysts are often limited by the data that's in the warehouse already. Instead of asking a data engineer or developer to do something so that more data gets pushed to the warehouse, I wanted to be able to pull it in myself.

Because of that, I'm starting an open source project, WebCrepe, to empower analysts to pull data directly into their databases using SQL. The idea is that we pair a database extension with a web app to enable searching the internet and pulling in structured data. It's really early right now. I have a docker-compose file you can use to spin up a postgres database and the backend. I still need to write some better documentation on how to write queries but it's basically using the advanced google search language.

I'm interested in analytics folks that have use cases I can build out and engineers interested in working on it. If there is any interest then I'll write up better docs and build more functionality.

I'm building a postgres extension that allows you to do web searches using a SQL query. The idea is to be able to pull in data from the web with some structure (which you define using custom scrapers) on demand.

Right now I have a proof of concept that's pretty simple. It's a multicorn extension that calls to a FastAPI backend. I have it all running using docker-compose.

I'm open to working with people that want to use it, or people that want to build it. I don't have any real plans to open source it or commercialize it. It's just a little side project I think is neat. I'm open to any ideas or use cases you might have.

Send me an email (in profile) or dm. Looking forward to it!

elovee (https://elovee.com) | ML Engineer, Data Scientist, Full-Stack Software Engineer | North America | Remote | Full Time

We are elovee, a healthcare startup focused on developing A.I. based technology to improve day-to-day care for seniors. We're building a voice user interface for seniors living with dementia. Our mission is to solve loneliness and isolation for seniors.

Roles we're hiring for - ML Engineer/ Data Scientist. - Full Stack Engineer

Why you want to work with us - We are small. You get to help set the culture and direction - Cutting edge technology. We are pushing SoTA Speech-to-text, conversation modeling, text-to-speech models. Tuning where needed and building what we have to.

What we are looking for - Experienced engineers that can take requirements and build products. - A strong sense of ownership. - Empathetic, team oriented teammates. - Connection to our mission

Please connect at careers@elovee.com or reach out with a DM.

Be direct, honest, and compassionate. They are good at something, but not the thing you need most now. Explain that to them, let them go, look for someone else.

Don’t forget that this person is a part of your network and always will be. They might be a good fit later. Someone they know might be a good fit.

This is kind of an odd question. It only indicates how old they were when you were lucky enough to meet them. Presumably, that person has been gifted before and would be gifted after as well, not just when they were XX years old.

I would be interested to know the trajectory these engineers had as they aged. Were they average until they got some experience under their belt? Genius from day one, but productivity improved? How did they develop over time and what were the predictors of greatness?

Yes. You should go. The job market is so much better. You will probably have a choice of high paying, interesting offers when you choose to leave your postdoc.

I did a postdoc in computational chemistry/condensed matter physics at LLNL and transitioned into industry in SF five years ago. I worked at two high growth startups and was promoted a few times. I got very lucky in picking good companies and got in as an early employee (around 10 each time). I gained a world of knowlege and experience in just a few years. Now I take that experience with me wherever I go. I would do it again in a heartbeat.

The options you have from Stanford will likely be as good or better than mine were. Best of luck!

This is a really nice summary of some of the technical components required. You also need to know how to do different kinds of analysis to answer different kinds of questions. A few more things:

0. Scientific method - probably true for all domains. Not really a kind of analysis, more an approach to doing analysis.

1. Cohort analysis - used in aquisition and retention analysis.

2. Model building - used in all kinds of financial analysis.

3. A/B/... testing - determining the difference between 2 or more populations.

4. Exploratory - understanding the relationships in your data to develop intuition about it.

There are plenty of analysis techniques in use. You can learn more about these and others if you survey blogs and other literature. One that I find interesting is Tom Tunguz. He has a particular theme, but his analysis is very good. The methods and way of thought are transferrable. http://tomtunguz.com/

Your email and post/comment history seem quite cogent. You still have the ability to write clearly and effectively. Please consider your options. Reach out to someone locally for help/contact. I would be surprised if there weren’t something you could work productively towards.

As long as you're having fun, learning new things, and not missing out on a great opportunity it makes sense to stay. That being said, depending on your job market, it doesn't hurt to look around. You don't have a hard decision to make until you have another job offer on the table.

So you want a LSM for inserts and the DNN for reads? Seems OK. You still have to update/retrain the DNN after an insert into a larger layer, which will be expensive. So you’d probably get high latency at the 99% (or some high number).

As it says in the paper, this might be useful for data warehouses. But, it’s not coming to postgres anytime soon. Index updates on the order of seconds to minutes would be too much for a transactional db.

There is also the cold start problem. How do you start to lay out the data on disk as you begin inserting it? Do you have a pre-trained net and use it at first (inserting where the net thinks the data should be)? The strategy probably differs by index type.

I've built a CNN before and that's not my understanding of it. The high frequency noise changes the output of the first layer on the CNN. This is what gets pooled as you go deeper into the net. Coarse graining is like getting rid of the weights you have for the first layer and replacing them with something uniform (average the smallest details together).

Adding high frequency noise "fools" ML but not the human eye. It feels like this is a general failure in regularization schemes.

Why not try training multiple models on different levels of coarse grained data? Evaluate the image on all of them. Plot the class probability as a function of coarse graining. Ideally its some smooth function. If it's not, there may be something adversarial (or bad training) going on.

Narvar | San Bruno, CA | Full-Time | ONSITE | VISA

We have a unique blend of data including logistics and shipment data from the carriers (UPS, Fedex, USPS, etc.), web analytics from customer activity, and product and returns data from our retailers. We are looking for data scientists, data engineers, and back end engineers to help build out new features and products using these great data sources.

You can learn more about us and apply through angel: https://angel.co/narvar/jobs or reach out directly to me jeremy at narvar

Narvar | San Bruno, CA | On Site | Full Time | Visa

We are helping e-commerce retailers provide an excellent post purchase experience for their consumers. Our current product consists of 3 parts: tracking, returns, and analytics. Our tracking product is a retailer branded experience that allows their customers to track their packages and shipments. Returns provides an easy to use returns product for customers to exchange or return products. Analytics provides insights into the post purchase experience such as shipping performance, customer satisfaction, and marketing asset performance.

We have significant traction in the market and recently announced our series A (google us!). We are looking to accelerate our current feature development as well as build entirely new products.

We are looking to grow across our organization including

    - Front-end engineer 

    - Back-end engineer

    - QA and test engineer

    - DevOps

    - Technical account management

    - Most other business functions (sales, product, finance, ...)
Our engineering team consists of about 15 people including 2 QA, 2 data scientists, 5 front end and 5 back end engineers. Our back end stack is Java and AWS with some go, and python. Our front end is javascript using some bootstrap, freemarker, jquery, and less.js.

You can find the detailed job postings here: https://angel.co/narvar/jobs

Apply online, or send your CV to jeremy at narvar

Narvar - San Mateo, Full Time, Onsite, VISA Hiring for front and back end engineering and data science.

http://angel.co/narvar/jobs

Narvar is a fast-growing cloud solutions company poised to change and disrupt how businesses handle their Supply Chain Management and customer post purchase experience. We use open APIs, SaaS technologies and are taking a smart, practical and data driven approach to supply chain. We are a well funded startup with several marquee customers. With companies of every size relying on our cloud solutions, Narvar thrives on innovation and succeeds with talented and committed individuals and the best customer service.

Data Scientist - Full Time - Onsite: We are looking for a self-motivated entrepreneurial data scientist with interests in between engineering, statistics, and product. We have data products for you to help design, build, and deploy including recommender systems, natural language processing, an A/B testing platform, and numerous predictive analytics models. You will be joining a small team and will be able to make an immediate impact. Are you an expert in one domain and want to learn another? Do you own one piece of the data science stack and want to master another? Let's do it!

Data Scientist - recent grad - Onsite: You will be provided with mentorship and given a choice of problems: starting from one-off descriptive statistics, to developing predictive analytics, to developing production grade, high-volume machine learning APIs.

Front-end Engineer Onsite: You would be working with the design and development team to constantly create and improve the experience for end consumer while supporting the product team on behalf of our retail clients.

Full-stack Engineer Onsite: You will be working with every aspect of the product, to develop the experience for our clients and the end consumers.

Feel free to email me (lead data scientist) jeremy at

Narvar - San Mateo, Full Time, Onsite, VISA

Data Scientist - https://angel.co/narvar/jobs/47508-data-scientist

Narvar is a fast-growing cloud solutions company poised to change and disrupt how businesses handle their Supply Chain Management and customer post purchase experience. We use open APIs, SaaS technologies and are taking a smart, practical and data driven approach to supply chain. We are a well funded startup with several marquee customers. With companies of every size relying on our cloud solutions, Narvar thrives on innovation and succeeds with talented and committed individuals and the best customer service.

The position: We are looking for a self-motivated entrepreneurial data scientist with interests in between engineering, statistics, and product. We have data products for you to help design, build, and deploy including recommender systems, natural language processing, an A/B testing platform, and numerous predictive analytics models. You will be joining a small team and will be able to make an immediate impact. Are you an expert in one domain and want to learn another? Do you own one piece of the data science stack and want to master another? Let's do it!

About the data team: We maintain a set of ML APIs using a microservice architecture. Our tech stack is mostly python code deployed in docker containers using Amazon web services where ever possible. Our data group leans towards an agile methodology for iteration on existing services. We cut code and deploy on a weekly basis. For new products and services we plan a MVP and then get to work. We work inside the engineering organization, closely with product, and provide support for all business units. We work on both internal and customer facing solutions.

Narvar - www.narvar.com - San Mateo (Silicon Valley)

Hiring for front and back end engineering and data science. Jobs posted to angel.co/narvar as well.

Narvar is the complete supply chain management platform that’s helping the world’s best brands improve the customer experience. From consideration to fulfillment, and beyond, our solutions deliver world-class, data-driven experiences to better serve your customers and transform your business.

We work towards improving customer experiences and maximizing customer lifetime value for businesses through a smart, engaging, and analytics-driven approach to supply chain using open APIs and SaaS technologies.

Full-stack Engineer ONSITE: You will be working with every aspect of the product, to develop the experience for our clients and the end consumers.

Front-end Engineer ONSITE: You would be working with the design and development team to constantly create and improve the experience for end consumer while supporting the product team on behalf of our retail clients.

Data Scientist INTERN ONSITE: You will be provided with mentorship and given a choice of problems: starting from one-off descriptive statistics, to developing predictive analytics, to developing production grade, high-volume machine learning APIs.

Feel free to email me (lead data scientist) jeremy at