SyntaxHighlighter

Tuesday, November 13, 2018

Performance Matters

Performance matters. And when we develop code we want it to be lean and efficient. Those algorithm courses we took in college and the study of Big O notation surely lets us write performant software. Algorithms that are fast, using the minimal amount of memory are sure to wow our stakeholders!

[Those of you working in the embedded space can probably skip this entire post. Performance does matter where CPU, GPU and RAM are limited. I cut my teeth for years programming towards and recording Technical Performance Measures on avionics display software. Every millisecond counted and when our load time requirement of "3 seconds" failed because of "3012 ms" - anything is on the table to increase performance. If performance is the requirement, it matters everywhere, and you should care all the time! ]

When does performance matter?

Unless performance is part of your requirements, it doesn't.

Most game and app developers, or those using modern web technologies (front end or server-side), well, performance does not matter that much.

At least, it does not matter that much right now.

When developing a modern system, do not focus on the optimal performance of your code in the first pass. Yes, you might be able to hand code a more efficient sorter than .sort() and you might save a few bytes from your network packets if you minify/uglify your payloads yourself. You might even move an algorithm from O(n^2) to O(n log n) in only a few hours.

The thing is, it just does not matter right now. Until your code is deployable in production, time spent optimizing and tweaking for the sake of algorithm performance is wasted time. Get it working. Make sure it is robust, passes your tests and can be deployed. That is the only time to go back and optimize. Focus on writing readable, maintainable, testable, deployable code first.

Does performance ever matter

For most projects, it probably doesn't. It could be a good exercise to code some whiz-bang algorithm or try to reduce the number of statements or calls needed in a function. Once your software is working and deployed is the time to know if performance matters. Why spend time optimizing something that might change? Or a feature that gets cut?

Consider these things as you decide if might matter:

Human-speed vs computer speed

Computers are really fast. Humans are not. If you are optimizing for something that will be noticed at human speed, don't. Why worry about a list of 100, 1000, or 10,000 items? You CPU will burn through those operations whether sorting or mapping or whatever. If the function you are performing is part of a user interaction, they will not notice the time savings between 157ms and 194ms. And a computer will do a lot in those 37ms. 

How often does this happen

Some functions and data flows happen often and concurrently. Speeding up those could have noticeable gains, even to us slow humans. However, some flows happen very irregularly. We have a special admin action that will gather a certain report. The report is costly to run, but is only requested maybe weekly. If all our concurrent users asked for this report every few seconds, our system would come to a halt! Time to optimize! As it is, the time needed to optimize is not worth the single, transient blip. Let's spend our effort on more worthwhile User Stories. 

Just scale up

Hardware is cheap. Can your current performance metric just be solved by buying a bigger machine? If your DB layer is running slow, is it worth your time to profile and optimize all your SQL statements and maybe rearrange your data tables? Or could you just buy a bigger server? Same for CPU, RAM, Network, etc. Do not sweat these performance decisions too early (after your initial planning) because sometimes a bigger machine will fix a lot of problems. In today's cloud world, a push of a button and little more yearly expense could save hundreds of hours of Dev/QA time. 

How to know if performance matters

Ultimately, you'll never know if performance matters if you are not measuring it. Performance metrics are absolutely necessary to 1) know where to focus and 2) know if it worked! Anyone that tells you "that's slow" or code reviews "that's not efficient" has absolutely no data to back that up! What is "slow"? What is "efficient"? Is there a suggestion that will satisfy this arbitrary feeling of slow? Until you can measure performance, you cannot know what is not performant or if you got better!

For a modern web system, you will see performance losses at the system (and typically network) level more than your code. Making 10 HTTP calls in series? You might notice that. Consider how those can be reduced or run in parallel. Is your DB more efficient with multiple JOINs or separate queries? What user actions are used the most? What are the slow ones? What is most efficient to make stakeholders happy? 

Start measuring and measure as much as you can. Then you'll really know where performance matters.


Using Puppeteer in AWS Node ElasticBeanstalk

Our company project has to send out reports and I like the flow of generating those reports from a web page to PDF conversion. We can reuse the tools and skills we use on our web analytics dashboards to develop the reports, with the added benefit that the reports look exactly like the web pages.

I first saw this flow from PhantomJS, the scriptable, headless browser. Load a web page into Phantom, say "gimme a PDF of that web page", re-configure settings 50 times to get it to look better, and boom!, you have a PDF of that web page. This is what we were using in our NodeJS server until reading this note and running into this issue. So we switched to Puppeteer, a similar project that was made for NodeJS and backed up by Google Chrome team.

Our servers deploy into AWS ElasticBeanstalk, their Node Environment. We haven't the need for containers or ECS yet, though did need a trick to get Puppeteer working in project. Things worked locally, but certain required things were not on the AWS Linux machines we were using, and getting those libraries installed on Beanstalk was a bit of a hassle.

A lot came from this StackOverflow question/answer (which itself came from somewhere else). Ultimately, through some trial-and-error, I created an ebextension config file to allow Chromium to install when Puppeteer is installed for the Node project.

File is below, and locally at .ebextensions/chromiumpackages.config. These run in order, first the yum packages and then the mysterious rpm commands. Note the use of --replacepkgs, otherwise your script will run the first time and fail subsequent times because the packages are already there. Yes, I guess technically we are downloading files and try/failing rather than checking if they exist, but it sure does keep the file simple!

And that's it. This config file runs when beanstalk makes a new instance, and then Chromium installs cleanly when my NodeJS project installs puppeteer. Yay!

Monday, January 8, 2018

NGINX HTTPS Redirect on AWS Elastic Beanstalk

We run a NodeJS app on AWS Elastic Beanstalk. Their node environment works well for us and on the machine is NGINX running on port 8080, and our node app running at port 8081. Some proxy table forwards port 80 to 8080, so NGINX receives all the traffic.

The SSL certificate is managed by AWS Certificate Manager, and is configured at the Load Balancer. That load balancer is serving traffic on port 80 (HTTP) and 443 (HTTPS), but always to port 80 on the beanstalk instance(s). It's very good to not have to worry about SSL in NGINX or node, and I could access our pages via HTTP or HTTPS.

However, I always wanted to use HTTPS. How could I automatically forward? Changing the Load Balancer to different ports (like 443 -> 80 to 443 -> 443) would have required some deep system changes on the beanstalk instances, which is handled by some complex beanstalk scripts. Examples are out there, but it all seemed overly complex and then I was tying our system to some random internet script that I did not write or understand.

Thankfully the solution was much simpler.

This is in my NGINX setup:

  server {
    listen 8080;

    if ($http_x_forwarded_proto = "http") {
      return 301 https://$host$request_uri;
    }

    // lots of other stuff
  }

That's all it took, and I understand it! If the protocol is "http", redirect for "https" and keep everything else (host and URI) the same. Mostly this will only affect the first load of the page, after that all future requests will remain HTTPS anyway.

(Only downside is sometimes I forget the "s" when pasting a `curl` command, and the response doesn't follow the redirect, responding with 301.)

Wednesday, December 13, 2017

SSH EC2 to EC2 and Security Groups

I ran into the case where my shiny new Elastic Beanstalk instances wanted to talk to some older services running on standard EC2 instances, and the method to do that was via SSH.

Problem


I use SSH often to manage that older service and connect with an Elastic IP (which is basically a static IP) as the "hostname". Trying this same approach worked in dev (our office has a static IP and it's just like my terminal), but failed when deployed to Elastic Beanstalk.

Managing the Security Group (SG) of the older service required a new access rule for port 22:
Using the SG of Elastic Beanstalk failed to connect.
Using the private IP of of our VPC (e.g. 172.31.0.0) failed to connect.
Using the public IP of the Elastic Beanstalk worked!

However, this was problematic because the IP of my elastic beanstalk could change (our staging system is a single instance but production is a rolling cluster). Editing the SG manually would be dumb and writing a script to check public IPs in the Beanstalk after each deploy sounded hard.

Easy Solution


Instead of referencing the static IP address XXX.XX.XXX.XXX when connecting via SSH, I used the public DNS, which contains the elastic/static IP address anyway (e.g. ec2-XXX-XX-XXX-XXX.compute-1.amazonaws.com, so shouldn't change on me).  This allowed EC2 to resolve internal IP addresses, and thus the Security Group rule on the older EC2 instanced worked for another security group instead of public IP. 

I panicked and asked on AWS Forums and StackoverFlow as well, and answered my own question at each. 

Tuesday, August 22, 2017

Why we chose actionhero?

While there are many frameworks for NodeJS servers, I have been using a framework for building our applications for the past few years and wanted to sum up those thoughts here.

TL:DR

actionhero


Production API Server

I've read a little about Jade and EJS as rendering engines for Node servers, namely Express, but never wanted to jump into having my front-end and back-end running on the same server. I had brief thoughts on the performance implications of having my single-threaded Node process serving and rendering web pages (over nginx or other proxies), and there were other reasons along those lines, though the scale of these projects were not going to be web-huge anyway. Mainly, I just did not want to tie my web page's HTML with server models, forever locking both together. (Having inherited a Java/JSP system, pulling the two apart in any language would leave some scars!)

Actionhero is almost purely for an API. Yes, it can serve static files, but to me it was built for data processing and message passing. I had flexibility in my client technologies because they simply spoke to an API. There are many features that are production ready and tested. This framework can be used for small prototypes to full-scale production systems.

Organization

My first Node project was built on top of Restify which was a fine choice for our data system (different team was building the front-end, another benefit of an API server). The concepts of Restify were solid, but I ended up developing my own system of organizing all my scripts. While that taught me a lot about Node's `require()` and the loading order of a Node app, it was a pain! Making sure all my end-points got loaded with clever directory layouts. Time spent on avoiding mixing the response handling and data retrieval. In the end it worked but there was time lost building, essentially, a framework for our app from scratch!

Actionhero has a folder structure for components of the system: actions, initializers, configs, and tasks. Each component type has a purpose, code layout and sometimes a load order. There is a flow to how the server works - where "actions" are the focus. Starting a new project with actionhero, or jumping into another actionhero project, has everything laid out the same way and I do not need to worry about it. Another win! Personal favorite concept of organizing code with actionhero is thin Actions and fat Initializers - the concept works so well in practice.

Features Included

Many of the other web server frameworks pride themselves in being "light" (Restify, Express) which is good, and then I'm spending lots of time choosing logging frameworks, middleware, websockets, config environments, scaling, etc. Flexibility is great, but the outcome of my choice was probably not vital, plus would have cost the time of research and integration.

Actionhero has some opinions and I'm content at the extent they are chosen. When a html app needed events pushed to it, actionhero already had a chat system and WebSockets. When a client did not want WebSockets in their Unity3D app, actionhero already handled every HTTP action via a TCP server. When I needed a dev, test and production environment for our system, a simple config controlled each one. A large list of common features to any production Node system are "built in", and while certain pieces are very opinionated, there is a lot of flexibility for everything else.

Documentation and Community

Final thing worth mentioning were the docs, which I've linked to a few times already. They are well written and comprehensive . . . most any answer I looked for could be found there first. And since just enought is included with the framework, and my code stays organized, I find myself writing more useful code than figuring out how to write the code.

Also, the community of actionhero is top-notch. This is due to actionhero's creator and top-contributor, Evan Tahler, who runs a great project. It's tested, documented, and evolving - and has been the past few years. The Slack channel is particularly useful these days, even over Stackoverflow or Github issues.

Shortcomings

Just to list these, these are my two short comings with actionhero worth mentioning.

  1. Redis is required. At least to utilize the chat or scalability (different server-nodes communicating) features. While a caching database isn't a terrible idea anyway, it's one more piece in the cloud infrastructure (usually one of more costly pieces) and there is not a choice if you'll only use some other similar DB.  For me, Redis is awesome anyway and until you need nodes to scale, it's not even necessary for a single-server-node (or server-node independent system).
  2. Ecosystem is small. When compared to Express, probably the most popular Node web server framework, there are far fewer plugins, blogs, examples, stackoverflow questions, and everything else for actionhero. Just not as many people use it, though due to documentation and community I have not found this to be a problem. Also, due to code organization, I know how to wire some npm package into the framework, and can use that package's documentation to figure out any issue there.


If this is your first node project ever, follow along a blog for Express to make a HTTP endpoint respond "Hello World." If you want to make a full-sized web application server, and be ready for all those production issues and feature growth you aren't even thinking about right now, get started with actionhero. That's my go-to choice while working with NodeJS.



Thursday, January 12, 2017

Upgrading postgresql tools on Amazon Linux

Tools like pg_dump, pg_restore and psql.

I didn't find much help when trying to do this, so thought I would write it up here.

Our AWS RDS was Postgres 9.6.1 but the postgres tools on the EC2 instance from the default yum repos was 9.5.4. Can't `pg_dump` with the minor version mis-match!

These were my commands (lines starting with // are like commands, don't run those : )

// Start at yum.postgresql.org 
// https://yum.postgresql.org/repopackages.php#pg96
// Reading through docs there, need epel repos enabled
sudo yum-config-manager --enable epel
// Now there are more postgresql packages listed, but not 9.6!
yum list postgresql*
// Need to install the matching rpm file with yum. Copy the link.
sudo yum install pgdg-ami201503-96-9.6-2.noarch.rpm 
// Yep, can see the packages for 9.6
yum list postgresql*
sudo yum install postgresql96.x86_64
// there was an error. Uninstall the 9.2 package (how did that even get there?)
sudo yum erase postgresql92
sudo yum install postgresql96.x86_64
// All tools now report 9.6.1
pg_dump --version
pg_restore --version
psql --version

Friday, November 11, 2016

Everyone Has Good Ideas

Everyone has good ideas. At least someone thought enough to give the idea. And while some ideas are really not that good in the end, you do not know until you hear them. When building a software product, ideas come from everywhere.

The business-side will have some ideas and you know they will be good ones, because they are paying for everything! Amazing how ideas sound a little different from the person (with)holding the checkbook. Developers will have good ideas because they are creative and just finished reading medium.com, and are eager to try new framework/library/database because surely it will be amazing and solve all the things. Other people in the company, or even individual clients, will have ideas now and again, sometimes in a very niche part of the software. And those ideas are important to them, more important than your next major or minor release!

What to do?!! Blowing people off and just blazing ahead with your plan might turn out well, and you'll prove to everyone how right you were! But at best some people are not happy, with you or your processes. At worst, you stop receiving new ideas at all. (Let's not even consider ignoring others ideas and still sinking . . . uh oh for you!)

When it comes to receiving ideas, two things need to happen:

People need to feel heard
People need feedback

Feeling Heard

Like I mentioned before, any idea someone brings up is important to them. If they do not feel heard or valued when giving their idea, they might become alienated to the team or just bitter. Could happen differently for different people, but the basic principle is to treat people well! We want good ideas, and we want people to give them.

Feeling heard is just a quick discussion, and it is also helpful in clarifying the idea.

  • "What were you thinking?" 
  • "Why is this important?" 
  • "Any timeframe the idea?" 
  • Maybe there's a quick drawing on a whiteboard. 

If you can repeat back the idea after you understand it, that person will feel valued.

All this will go into your backlog of choice, whether its a spreadsheet, sticky notes in your dev's cube, or some tool that team uses. Ready to evaluate and enter planning (or cold storage : )

If you hear people's ideas and they have no way to get any feedback, their ideas just go into a blackhole and they have no clue if the idea is something to prepare for. We contracted on a team once where the Tech Lead wanted the business team to leave the dev team alone, constructed a wall, and ideas were tossed over (through him). Without any feedback from their meetings, the sales team were pitching one set of ideas to clients, the dev team was building another set of ideas, and everyone only found out how far apart they were at Milestone Releases! Contrast that with a company I've spoken with that has great intra-team communications, and ideas are flowing!


Giving Feedback

Often called closing-the-loop, any idea needs a resolution.

  • Is it happening this sprint? 
  • Is it up for future review? 
  • "Yes, though wait until Milestone 2 is over." 
  • Skip it, because...


It does not matter whether this feedback is personally communicated or can be looked up at another time. It does matter why the resolution turned out the way it did.

People need to know what was good about the idea, and what was not. Maybe the idea was good, just not right now. Maybe the impact would not be worth the investment. Whatever the reason, it communicates the goals and schedules of the team, and will help shape better ideas in the future.

If people's ideas get feedback, but they are not heard, the people do not feel valued. Ever hear things like this?

  • "Why bother giving ideas, they just get rejected?" 
  • "Those tech guys don't know anything, this idea is necessary and will save the company!" 
  • "Yeah, I saw the demo. That was my idea, ya know?"



People that feel valued will contribute more, will stay with the company, will think of real ways to improve because they genuinely care. Good people are the greatest asset, and must be valued (of course, a principle in life and not just business.)

Now Do It


The process your team implements can be adjusted, but both parts are necessary.

The existing meetings (Planning, Retros, "weekly team", etc) you already have are a good place to touch on these things. Keep using whatever tool you're using to manage tickets. With an attentive team lead (scrum master, whatever), even "being heard" can be accomplished remotely through the ticketing system; though starting that way may not have the right affect until people can trust the process. Even give rewards for good ideas. Whatever it takes for your specific team.

Everyone has good ideas, so work hard to find and keep the good ones, and ensure the next ones are brought in.