About this talk
This talk discusses the real-world experience of Encore Healthcare in achieving SOC 2 Type 2 and HITRUST certification, specifically focusing on compliance in software development and population health management. The speaker, Matthew Conton, who has extensive experience in the Drupal community, emphasizes the importance of continuous compliance practices rather than a one-time checklist approach for audits. He explains the various frameworks involved, such as SOC 2 and HITRUST, and highlights the significance of robust documentation, logging, and the implementation of best practices in security and system management. The presentation covers tools and technologies used, including AWS for hosting and Drupal for application development, while also addressing common challenges faced during the certification process. Ultimately, the session aims to illustrate building a compliance-focused corporate culture that integrates with engineering processes and fosters continuous improvement.
Full transcript
Okie dokie. Well, thank you so much for coming to this presentation where we're going to share real world experiences uh from successfully achieving SOCK 2 type 2 and high trust certification at our company Encore Healthcare. Um I am Matthew Conton. Uh I'm am the VP of software development at Encore and I've been in the Drupal uh community for about 18 years now. Uh, Encore Healthcare is a
software developer and population health company that focuses on providing better outcomes for COPD and chronic respiratory patients while reducing health care costs for the health systems and the payers. So, this session is not going to be a straightup checklist on you know do these things and you will be compliant. We're going to talk about building processes you know building the engineering tools required uh so that the
audits are easier. Um compliance is going to be a continuous effort uh not just something you start and end uh when it needs to get done. Uh if you are, you know, scrambling before every audit, it means your processes aren't working. You're just cramming for the test at the last minute. Every section of this talk is going to help map to one side of this table. And
by the end, uh, we hope that, uh, you'll see how the practices move from the left's column over to the right and you'll be able to deliver, uh, auditable materials continuously. Uh, I want to talk about some of the uh, glossery of terms we'll be looking at. If you're not familiar with it, SOCK 2 is a framework developed by the American Institute of Certified Public Accountants to
evaluate how service providers manage customer data across five trusted service criteria. There are two types of SOCK 2 audits. The first one is a point in time audit and it's a self audit you perform by yourself. And then the second one is a period audit typically over 6 to 12 months. So type one is do you have controls designed correctly today and type two is did those
controls operate effectively over a period of time. High trust which is what we're going to primarily be talking about our experience with is a comprehensive prescriptive security and privacy framework built around healthcare regulations HIPPA high-tech ISO PCI. You can think of high trust as SOCK 2 on hard mode. uh more controls, more documentation, more rigor. Uh and like I said, I I'm very ambitious with this uh
this presentation. There's lots of slides. You'll probably see me skip through some, but they will definitely all be available. Um there are other frameworks you might have heard about where these practices will also apply. Um there's the ISO frameworks, the HIPPA, PCI, DSS frameworks, Fed ramp is a security framework uh where a lot of these practices will come in handy. Uh and lastly before we dig into
the practices, I want to talk about our product Nexus that uh we use to provide compliance uh for these patients. Um it is a multi-tenent clinical SAS platform. It's a fully decoupled application. Our back end is in Drupal. Our front end is in remix and we are hosted in uh AWS. Um so we're using all the typical AWS uh systems VPC isolations roles encryption secret managers everything
you would expect. Everything we do in AWS you can do in G uh CP you can do in Azour you can do on more or less Lenode and Digital Ocean these days. Uh all of the uh topics apply. Um, let's talk about what's in scope for a certification like this. What do the auditors actually want to test? What are they going to review? For us, it's our
critical production systems, especially systems that are going to hold uh PII, personally identified information and PI, PHI, personal health information. Um, our production infrastructure was the only thing in scope as far as our application uh the front end and the back end. But at the same time since our uh work computers and phones could have PHI on it, we had to have endpoint management on those devices
as well. So they become in scope as well. Um the people that have access to these systems are also in scope. Uh engineers, my devops team, the support team, the account managers, all of their admin accounts or accounts that have access to PHI become a part of the scope. uh things that are not going to be in scope is for us was anything that's not production. If
it wasn't uh hitting PII or PHI that um had patient data in it, then it wasn't going to be a part of the scope. Uh the corporate office facility we have, we're distributed team, but our offices in Tennessee um that was out of scope. Uh you know, we have locks on the door, but we don't have any servers there and our laptops, our machines are in scope,
so the physical building wasn't. uh most of our internal business apps are not going to be in scope unless they are explicitly being used for PHI. So the exception for us is uh we have enterprise Slack with a BAA signed with them so that you know we can have uh HIPPA compliant uh communications in Slack. Um none of our customer environments are in scope. That's their responsibility.
uh and none of our nonprod CI/CD uh resources are in scope because nothing production data uh goes across there. So let's talk about the certification process and again this is going to be the high trust journey. This is uh a very strict workflow and process we went through. Other systems may be worse. I'm not as familiar with Fed Ramp as I am with this. Um so 2
is a little bit easier but that's still relative. Um there's a lot of work to it. Uh there was a readiness assessment that had 25 requests of controls that we had to fill out and we did that over a couple of months and we thought, "Oh wow, this is easy." You know, we got this covered. And then the actual uh assessment came out with 348 requests and
we were, "Oh, wow." Uh it was a lot. Um the uh you're going to want to have a consultant unless you have a dedicated security person on your team with experience uh with this type of certification. You're going to want to find a partner uh that can help guide you through all these controls uh help guide you through uh all of the requests uh and be the
in between with the actual auditors. Um we used a company called Thorpass uh and they did great. um review these requests with them as early as possible. There's lots of ambiguous controls like verify that your input matches you know the expected output. Uh and you know that could be interpreted a dozen different ways. You have to see how does that apply to your system what you're doing.
Um ask for evidence uh for these examples for those ambiguous ones. What exact what's going to satisfy this? What are you looking for? Um plan for a long audit. Again for high trust for us um it was about 18 months total from start to finish. The first two to four weeks was just preparing scoping seeing what was uh and was not going to be included. Then we
did that readiness assessment uh and that took about two to three months to fill in the gaps see you know what needed to be completed before. Once all of that was done uh we did the actual evidence collection for the period of time that was about 6 to 12 months. This is the longest phase. This is where you can implement and prove your controls. Uh and you
have to show that you know you're implementing the uh controls via evidence. There's lots of screenshots uh involved here. If anybody has a technique for taking a screenshot of your desktop with a clock visible, I'd really like to know that. Uh once all of that is done and submitted with your consultants, they're going to validate it and hand it off to the auditors. They may come back
to you and say, "Hey, can you give us a little more detail on this? I we didn't quite understand that." Um, but at that point, it's in their hands. And then for us, it goes to a final uh high trust QA uh certification process. I'm sorry, high trust QA process. And then once everything passed, we got our uh certificate. Um again, we took about 18 months from
start to finish. We had already um done our sock type one. We had an overlap with our sock 2 from when we were completing it type two with uh when we were starting high trust. So we had a lot of the uh practices and work already done. Um if you're brand new to this, you don't have a strong security posture, you know, expect up to uh you
know, two years to be fully prepared to do all of this. If you're coming to the table ready to go, then it can certainly be a lot faster. You know, a lot of the delays you're going to run into is when they're gathering the evidence. Um it it's not just a single point in time you did this thing on that day. You have to show that you
were adhering to policy procedures these controls throughout the entire period that was requested. Uh typically it's going to be about 90 days. Um we had a couple of vendors that we didn't have a BAA with and it was questionable if we needed that. So we had to get that resolved pretty early on. We had disaster recovery setup uh plans but we really weren't testing them. you know,
we had business continuity plans set up, but we didn't have a tabletop to say, "All right, our CEO got hit by a bus. Well, what do we do?" Um, you know, change management, asset inventory, a lot of the things that I'm going to talk about here in this presentation, um, if they're those are the delays, getting those set up if they're not done ahead of time, uh,
is what makes these things go longer. So, if you can build this into your, uh, corporate culture early, it makes it all easier later. Uh, and again, find a consultant to help with this, unless you got a uh, expert on your team. So, let's dig into the actual meat. The most important thing of this entire process is to document everything all the time. Um, if you do
this, you're more than halfway there. Uh, for high trust, this section covers 60 plus change management controls and 50 plus documentation controls across SOCK 2 and high trust frameworks. Without proper documentation, you can't prove the controls are operating effect uh effectively. Regardless of how secure your systems actually are, every decision that's made needs to be documented in some form of uh ticket system and it needs to
be approved by someone um before it can proceed. Even if you're not quite up to par on some of these other controls, if you're documenting why you're not up control uh up to par in mitigating it in some form, you can still get a high enough score in that control to pass the assessment. Store your documentation in a central auditable uh repository. You know, for tickets, really
any ticketed system will do. It doesn't matter. Um for other documents, it could be Google Docs, a wiki, Confluence, you know, something similar. Just pick something, stick to it, be consistent with it. Uh, you can even document uh in a codebase with markdown files. Um, just about every part of this presentation is going to be about what you document. Um, but the key here is start early.
Start as early as possible. Every little thing, um, create a ticket for it and then say, "Hey, you know, we did this thing." And get someone else's approval to, uh, to push it live. Another core tenant of uh the compliance journey is policies and procedures. There's a lot of them. These are the core of what the rest of what we're going to talk about uh where all
the rules come from. So start with templates, you know, Google them. You'll see a list uh that I have here when you get my slides. Um if you start with a consultant, they're typically going to have a set of templates for you to uh use and modify. uh you know it's it's not just tedious paperwork. It is a lot of uh tedious uh uh writing but um
it's the core rules of how the rest of your compliance system is going to behave. Um we have about 24 um procedures and and policies. Um I'm not going to go through each and every single one of these, but you can see there's a lot of fine detailed things um that matter. And once you document the rules uh like your password requirements or your you know website
privacy policies then you know how to uh implement those later. Next up is logging everything that happens all the time because if you don't log that it happened then it it didn't happen. Uh and a big part of logging is context. Uh you want to capture as detailed context and metadata as possible. Um for a multi-tenant system like Nexus we include tenant specific context such as the
uh account the organization the user identity and roles and any trace ids or you know additional context from wherever the log is coming from so that wherever these logs show up we have we can group them by or this is this tenant this is that tenant this is this situation. Um you really can't overlog when it comes to compliance. um where you going to store these logs.
If you can't over log them, start simple, you know, start with your DB log and when that becomes untenable, start dumping them to SIS log or to JSON files. Um we are starting to put logs or right now all of our logs go to data dog, but none of them have PHI in them. Um but we want to start logging uh some logs that have just a
little bit of PHI to help with uh resolving issues and uh raising exceptions to our end users. So uh we're working on using uh open search in AWS. It's cheap. There's tools readily available to take those Drupal logs and aggregate them and ship them or ship them to open search so that they can be aggregated. Um there's tons of platforms out there like New Relic and Data
Dog that can also uh aggregate these logs for you. Once you have those logs, you have to keep those logs. Um for SOCK one, it's for at least one year. For sock two, I'm sorry, for sock two, it's at least a year. And for HIPPA, um, and high trust, we have to keep these logs for six years. We keep them in a, uh, warm storage. And, uh,
after about a year, I think we dump them to uh, cold storage. Uh, Libby, can you come plug this in while I'm talking, please? And thank you. So what are we going to log? Uh definitely all the standard stuff with Drupal. Uh the standard logs we log anytime configuration is changing. Um we have a lot of cues that run. So automated jobs and background processes. We have
them dumping to the logs. Anything security uh relevant gets logged. Um in our system we don't delete anything. None of our entities get deleted. So we have a lot of revisions and those revisions act like logs of things that changed at a time that a specific person did something. On the AWS side, Cloudatch is aggregating a lot of these uh resources logs. Cloud trail is monitoring and
logging the you know user account uh usage in AWS and then guard duty uh is watching over everything on top of it. There's also logging you can do uh outside of uh the logger. Um we have a activity record that whenever a clinician is downloading uh some reports, it logs that they downloaded it, the date they did it. Um so that the next time uh they go
and see this particular patient, it can show the date that they uh downloaded that report. Or if there's some new data, it can show the download button again. when we're syncing data with third parties, we have the sync window entity that can record, you know, for a ventilator, this data that got downloaded from this time period. Um, so that is ways you can capture, you know, log
type of information and events. Once you have these logs, you can uh, you know, you want to set up alerts and intrusion detection on them, maintain user and especially admin activity logs. And then there's uh lots of tools you can use to set up automated uh detection for anomalies. If there's a lot of failed password attempts for a user with an admin role uh after you know
n number of times trigger an alert that someone maybe blocked the account. Observability uh is going to be continuous compliance here. Logging lets you know what happened. Observability tells you what uh that your controls are still working. This is the difference between the seasonal audit and preparing for continuous uh assurance. You know, you want to be able to map these things that you're logging into alerts and
controls. If there's, you know, failed login spikes, that's, you know, your access control and intrusion detection. If you're getting a lot of unauthorized API access, that's checking your lease privilege in boundary protection, those controls. So all of these things that we monitor can map to specific controls so that when the auditors say how do you know that this control is operating effectively you can have a dashboard
that says well we're monitoring this right now here are the configurations for each of these controls and for each of these monitors and how they apply to the controls. Data dog does a lot of this for us u but cloudatch and AWS config and guard duty can also cover a lot of this. Um I have this uh slide here about building compliance dashboards. You know if you're
using something like data dog um it can certainly help uh you know it's less helpful for when you're looking back at those 90 days but it's very helpful to make sure you are doing that continuous auditing uh of your controls. Um there's some Drupal modules that help in this uh department. Monologue is a fantastic one. It lets you just take all of that watchdog data and ship
it out in the different places. We're working on having it set up to send the specific um ph non-phi stuff to data dog and then the phi logs will get split off and put into open search. There's the admin audit trill uh which can track your admin and config changes. You know sis log is a really good one especially if you're in um containerized environments. You can
dump your logs of sys log and whatever is orchestrating your environment can capture all that up and ship it off wherever. Um, and if you're using something like a activity entity or audit entity, you know, search API is a great tool to index those events into a external system that can then be searched and audited on. So, user onboarding is about, you know, not just about users
on your site, but employees and contractors of your company. You need to document every time you hire somebody, whether it's a vendor, a contractor, anybody that's being granted any sort of access, have a ticket template for that. Have a checklist of all the steps that are required um to go through that process. Um access account provisioning. Uh auditors are really uh happy when we use SSO because
it's a single place that can enable or disable access and they all have multiffactor authentication built into them. Um, so try and use SSO as much as possible. Um, only provision access to, you know, what's necessary, not the junior devs don't need production access on day one. Um, make sure you got role-based access uh based off of people's job functions. Um, lots of uh places are going
to have legal and compliance requirements. So, background checks, NDA, the auditors want to know that those are happening. They want to know that somebody did it, checked it off in the ticket, and there's a document in secure storage somewhere that it happened. You're going to probably require some online trading. Uh with us, it was uh you know, how to handle PHI, where we can and can't use
it. You know, make sure the people know about social engineering attacks for the developers. You know, safe and secure uh coding practices. Um, you know, it it's some of them can be a little cheesy, like they were written in 1996. You know, don't pick up the USB stick and and put it into your computer. But, um, you know, some people need to hear that. Apparently, if you're
issuing them devices, make sure that they're enrolled in your endpoint management. Make sure multiffactor authentication is set up. Make sure you have security baselines um, with disk encryption and screen locks and antivirus. All of that should be controlled from a central location. Um, we provide evidence for this type of stuff by showing them the onboarding ticket, all the signed documents of completion and NDAs and background checks.
Uh, what access check was granted and what the approvals in the ticket that yep, this person has access to this, I approve that, I approve this. Training logs that got completed. Uh, any exceptions that were documented and approved. All right. What about when you offboard people? Well, you do everything we just talked about in reverse. Disable the accounts. Again, SSO makes that really easy um to turn
off a lot of stuff in one go. Revoke access in other places if necessary. Make sure you collect devices, remotely lock them if it's uh a situation where you you need to do that. Um remove them from password managers. Again, same thing in reverse. Um there are some Drupal modules that can help with uh identity and authentication. There's lots of modules that integrate with SSO um using
OIDC. Um you know if you can avoid using local accounts that's great. If you can't uh that's fine as well. We still have local accounts in our uh system but our we also support SSO for our customers. They can bring their own um SSO to log into Nexus. Um, civil oath is a great module for handling access control and scope into your sites and APIs. Uh, the
TFA module gives you two-factor authentication in Drupal. And the password policy is fairly critical for ensuring that you're matching your password policy. So once you have these users, you need to, you know, again, control what they have access to. We've talked about some of these already. least privileged. Everybody doesn't need access to everything. Make sure FF uh MFA is enabled. SSO um admin accounts was a interesting
one with high trust. They specifically wanted to make sure that our day-to-day accounts did not have elevated privileges that if we wanted to use pseudo or you know on my Windows computer go uh and elevate as an administrator that needed to be done in a completely separate account as a safety measure so that dayto-day I didn't accidentally uh you know trigger something or if there was something
infected on the computer it couldn't use my uh elevated account against me. um your apps and servers, you know, pretty uh straightforward stuff, but specific things they looked for was um default denying and networking. There should be no allow all rules in your security groups. Uh any of your sensitive data should be in a private VPC, restrict server access to uh secure entry points uh only. We
use tail scale uh and we SSO via Google to get to that tail scale access. Um, so we have pretty strict control over who's allowed to get into our servers and how they get there. Um, and that works out really well. Um, use a firewall, use a web application firewall. Um, we use AWS, but every major system has one. Uh, Cloudflare has one, is one, Azour has
one, Google has one. uh they're relatively inexpensive and they you know they do come in handy for some of those uh zero day uh attacks that get out there. Uh I think the Drupal association has one now um that you can get directly from them. You have all these access uh policies and procedures. you have to review them uh at you know quarterly uh at least and
what that means is we have a big list of all the administrator users and just a general roles and what they have access to at a user level and an admin level and uh four times a year we go through that list and say okay has anything changed for the CEO for the cso for the VP of this for the VP of that um and then below
that table we have a list of all of our services that we kind uh you know sync up there and see whether any new services added uh that need to we need to track access for um they expect to see that we had a meeting that a document was written that we went through and and uh said yep we checked all of this and everything looks good.
Um so for access policies uh and we'll probably start skipping these uh evidence provided uh slides here pretty soon. um you know everything we talked about just screenshots of the of the configurations and that the enrollment is set up in the endpoint management. Um you know this covers Drupal roles AWS IMA uh device management roles and network uh and WAF controls. Um there's Drupal modules that can
help with access control inside of Drupal. Um the group module is a fantastic one. Um, I was always a big fan of organic groups back in the day, but the group module is doing pretty good now. Um, we're actually have a custom entity solution uh that manages our our access control and we're shifting to use the group module now to reduce uh what we need to maintain
at the custom code level. Um, permissions by entity, role delegation, content access, you know, field permissions. There's tons of access control modules in Drupal that can help ensure that people are only seeing um and being able to have access to the operations that they're supposed to have access to. Keep your role small and purpose-built and tied to job functions. So, a secure SDLC or software development life
cycle uh is a huge chunk of how you can stay compliant and it all starts with documentation across the board. And a big part of it is going to be in your tickets. I'm going to talk about Jira tickets, but really any ticketing system uh you can use here. And the business and security impact uh has to be robustly documented uh in these tickets anytime you have
a change. What's the context, the background of the change? Why do we need this? What are the requirements? What's the impact across security, confidentiality, inte uh integrity, availability? Are there risks if we don't do it? Are we going to impact privacy? Does it involve PII or PHI? Are there ways we can minimize the data uh that this feature is going to uh touch? Are there regulatory impacts
for what we're about to do? Um what's the acceptance criteria for this particular ticket? Are there dependencies or constraints on something else? The auditors want to see that we've thought through all of our production changes uh that go live. You know, do you need to do this if you're just adding a single field? um we you know not really um but you need to do a light
mode version of it. Uh if you're implementing a you know epic or story level you definitely want to go through all of these uh Um you know I'll talk about it later but this is one area that AI helps a lot. Um, I have custom commands for generating requirements and specifications where I can tell it what I want to do and it knows I need to cover
all of these bases in my tickets. And so if I forget to give it content for one of these, it will say, "Hey, you know, you didn't tell me what the acceptance criteria was or is this impacting HIPPA?" But a lot of the times it can derive based off of what I told it whether it does or not or generate that criteria for me and I can
just tweak it later. Um, all of your releases that go live need to be documented. each individual change. It gets a ticket. Each entire release should have a ticket. We're going to release this whole thing. Uh and the status change of that ticket uh was a really key thing that the auditors looked at. They wanted to know that I had a release ready and I assigned it
to um my development manager and I said, you know, Carrie, there's 13 tickets in the release. You can review them here. We've already uh gone through the PR processes and tested and everything got approved here. This is everything going live. this needs approval to go live and she'll look it over and say, "Yep, because she's already done uh some stakeholder UAT as well." And they wanted to
see that it got assigned to her and then it got assigned back to someone else and that the status changed from ready to uh release and release approved. You want to document your infrastructure changes. If you need to go add a new server, add memcache, tweak this, tweak that in your infrastructure, create a ticket, and then make the change. make a ticket for everything. Uh, and this
is where, you know, AI can be really helpful in breaking up this tedious task. Um, the SDLC uh requires a separation of duties. Uh, you can't have one person that says, "All right, we need to do this feature." And then they develop that feature and then they create a pull request for that feature, then they merge their own pull request, and then they push it to the
production branch, and then they tag it, and then they deploy it. One person not allowed to do all of that. um there has to be someone else. I'm not saying that the developer can't literally push the code live. It just means that you know the developer needs to assign it to someone else for code review or there needs to be assigned to a stakeholder to do UAT
or maybe it's assigned uh at the release level to the uh developer manager that says this, you know, this has been done. It needs to be released. Do I have approval to do that? There has to be a um QA needs to match and UAT need to match your production. Uh they asked about down to the version level. Does our engine X version match what's on production?
You know, is the PHP version the same as on production? They want to know that you know what we've tested in QA and UAT is going to produce the same results in production. Um automated tests and security scans are going to be uh super critical here. They want to know that you know we have so many dependencies coming in through composer and npm that we're tracking that
we track it with a dependabot our code quality with PHP stand PHPCs PHPMD that all of those are being executed and run uh either manually by your developers or automatically in your CI/CD. Uh there was a one of those ambiguous controls was about uh validation the input and the output checking for accuracy completeness and validating the authenticity uh of the data and that was one that you
know we really went short what they were looking for here. Drupal, you know, entity system has obviously built in a ton of field validation that handles a lot of the import validation uh natively, but this is a great place for uh PHP unit tests where you can validate um your custom code at the unit level, the kernel level, the functional level uh for endto-end testing. Um there's
a great session on Drupal test traits, you know, combine that with DDEV. Uh and then and you know another place where AI shines, Claude is really good at making these tests. Um and it helps out a lot uh in helping uh you know feel comfortable with these uh changes and it gets another control done. Um high trust specifically wanted us to make sure with our releases they
were getting tagged and we were releasing archived. They wanted fully static deployments. They wanted production servers with no build tools. It shouldn't have git on it. It shouldn't have composer. It shouldn't have mpm. We should be deploying a static artifact every single time. Um, this is something that we don't do, but we have mitigate it in different ways through access control because we have a legacy system.
Uh, but if we were building something new, then we would end up doing this. And if again you're shipping in a containerized environment, that's building your images and shipping your images. Um, your pipeline is your audit trail when it comes to your evidence here. Uh, and it is a large part of it. Your pipeline is the evidence generator. Your PR requests prove your separation of duties and
your uh, peer review. Your CI runners that are doing your test and your linting, that's your code quality and input output validation. the security scans that run automatically whether it's dependabot composer audit npm audit that's your vulnerability management um branch protection rules uh prevent no direct pushes to production tagged release triggers deploys if you want to do fully automated deploys that way that's your trains change management
and traceability and then whatever is actually doing your deploying that's going to output some logs and and timestamps and know who did it um we're using uh Jenkins as our actual production deploy but it has a complete log of the pipeline every time I do it. Um so when the auditors say you know show me uh how this deploy happened we show the PR's uh logs output
we show the uh Jenkins release log outputs and a quick example of this is you know you have a feature branch linked to a ticket the code review is approved the CI passes it gets merged in the main you tag a release the deploy is triggered whether it's automatic or manual um and then the deploy log is capturing uh everything that happens it doesn't matter you know
if you deploy manually or automatic as long as it's logging that output and saving it somewhere. And this works the same whether it's, you know, GitHub actions, GitLab, Jenkins, CircleCI, TravisCCI, whatever tool you have, as long as it's not a person sshing into to production and running commands. Um, automate that deployment. It it could be a bash script that dumps to a file if that's what you
need to start with. Um, as long as you have that documented and reproducible. uh the pipeline controls that really matter is that branch protection um requiring some status checks uh if your tests are failing then that PR shouldn't pass uh signed commits for verified actors so you have verifiable information of who made that change um logged uh deploys uh immutable artifacts we talked about that part okay
change management and we're specifically talking about your infrastructure They want you to be using infrastructure as code for all of your infrastructure so that it can be uh documented and logged via your SDLC. Um there's tons of tools you can use for this. Terraform SST is what we use for our front end AWS CDK. Uh they're all valid options. It doesn't matter what you use. Ideally, you
want to use that. But there's times you can't. And it's when you inherit a seven-year-old legacy application that was built by four different developers and you know in a massive AWS environment single account. Um that's still okay as long as you have mitigating controls set up for if changes do need to be made, how are those changes made, who makes them, you know, are there separation of
duties there? Have it documented and make sure you adhere to it and have proof uh when those things change. Um any new infrastructure ideally should be infrastructures code. At that point you want to have all your uh devops in code. It could be a bash file. It could be anible. It could be Jenkins. It could be gitlab. Uh any of the devops that's happening with your production
instance deploying live maintenance. Whatever should all be in code via your SDLC. Don't make changes on prod directly. Uh you know obviously that's an easy thing to say and a hard thing to do sometimes. Um, this is not saying that you're not allowed in in the event of an emergency to do a hot fix, but if you do go make that hot fix, the moment you've put
out that fire, you then start the process of documenting it and following the SDLC from that point. Uh, environment separation is configuration management. You want your prod and non-prod stuff to be in separate uh, environments and and not just, you know, a separate database, not just a separate server. um at the organizational level with AWS you want separate accounts um so that there's no cross bleed accidentally
between you know production phi somehow getting access to stage or vice versa um no shared credentials no shared roles or resources between those environments um the we use AWS uh organization structure you know prod you could have non-prod environments we only really have one utility server for all of our non-prod stuff. Um, but if you had a Devon stage, you could do that. You could have an
account for all of your security logging and monitoring. Uh, and then the only shared services would be your CI/CD and networking type of artifacts. Um, you want to enforce guard rails between those environments. AWS has service control policies that can block unsafe changes. So that can prevent, you know, an administrator from going in and turning off guard duty. Um, even if they're an admin and have access
to everything in that account, the SCP will lock that down and prevent that from happening. Prevent them from turning off multiffactor authentication. Um, so leverage those. Uh, if you're not on AWS, that's fine. Again, every other system has some form of organization, group, project level that you can separate all of this stuff out into. um all of this stuff that's in those uh environments needs to be
tagged. Uh and what that just means is you know I have a server here. The classification has PHI on it. Who's the owner? What's the purpose of it? These tags are used so that they can say show me all the servers that are touching PHI and you can bring that up and then you can apply other controls over that to say okay do these servers you know
have the endpoint management are they logging correctly? are they doing XYZ? They help uh define and limit your scope. Um you know, everybody's is going to be slightly different. Define the different classifications. Tag them according to sensitivity, owner, and environment. Uh document the, you know, the handling requirements in your policies for each of those levels and restrict access based off of those classifications. This is a pretty
easy thing. Uh all of your stuff should be encrypted in one form or another. Uh at rest, we're encrypting our databases, your file storage, your volumes, your backups, they should all be just naturally encrypted. And you can enforce that with SCP. Uh any data in transit should be also encrypted, whether it's internally or externally. Uh document that encryption uh and your policies and how you're handling your
keys. Uh there's modules that help if we're talking about inside of Drupal. uh you know the encrypting key module uh or an entire framework and can work with uh thirdparty services for encrypting in and out of Drupal. Field encryption lets you encrypt individual fields. File encrypt lets you encrypt files. Um so there's lots of solutions available to you there. You got secrets you got to keep uh
and you should not keep them in git. Uh, so don't put any of your uh API keys, any of your secrets in git. There's a lot of modules out there that have config forms uh settings forms that say put your API key here and you put it in there and you hit save and you just cex and then you accidentally commit it in git. Um, thankfully uh
GitHub and other systems have tools that can detect that and say whoa in your PR looks like you committed a key. Um, let's rotate that and get it out of there. uh put those in environment variables or pass them in as uh environment variables and then use your settings.php to load them into the different places that they need to go dynamically. If you suspect that a key
has been compromised, just rotate it. Um use the dedicated secret management system. There's a dozen out there. It just depends on what uh environment uh and tools you're already using. We use a combination of AWS secret manager and one password. Um, yep. All right. Backups. Uh, you now have built a great system. You need to keep it backed up if the worst happens. Uh, backups protect against
data loss. They protect against corruption, ransomware, uh, you know, your junior dev getting pissed off and burning the system down. They are really necessary. uh if you uh core backup requirements all the data must be back backed up by uh specific frequencies your security framework and your policies are going to f define that you have a little bit of wiggle room there you got to protect them
with access control they need to be immutable um they need to be retained for a certain period of time we saw that very similar to the logs um what the auditors are going to be looking for are the formal documentation of your policies evidence your configuration of your backups. We use uh what's it called? AWS vault or backup something. I forget exactly what it's called, but you
know it shows what we're backing up, how long we have to keep them. We have it in I think it's compliance mode which automatically makes them immutable so that you know nobody can delete them or change them. Um and then uh everything is encrypted as well. Um, and then you set up via your uh, guard duty and and and Cloudatch if there's failures. Make sure there's alerts
set up on that so you can fix any issues. Um, what needs to be backed up is obviously all of your production data, file objects, images, volumes, your configuration and infrastructure is backed up in Git. Um, so that's great. Encryption keys and secrets. Uh, logs were backing up to S3 and then they get backed up somewhere else. Uh, and another one here that was unique was any
critical third-party SAS data that maybe you're integrating with another system. If their data is critical to you, you can't just trust that they are backing up their data if they have an issue. So, you need to back up whatever is critical uh in that um situation. A shared responsibility doesn't mean the vendor is just going to do everything for you. Um, we already talked about that. All
right. So, when you're backing up, there's these terms, RPO and RTO. RPO is your recovery point object. And the uh how much data loss is acceptable for you to be able to restore. Your RTO is how fast I'm sorry, the uh recovery time object, how fast must you recover after a failure. So, if there's a failure, you know, can you afford to lose an hour's worth of
data? um can you afford to be offline for four hours? Whatever those numbers are, put them in your policy and then um you know, make sure that you can actually meet that requirement because if you don't test your backups, you don't have backups. So, you need to make sure after you get those backups that you're able to actually test them and meet the requirements that you put
in your policy. Um when you're testing them, you want to be able to do full database restores. You want to be able to do point in time restores. Uh we needed to also be able to pinpoint individual records or file restores at specific times. Uh and then validate that RPO and RTO. Um test them at least annually, maybe more than that, maybe uh every time you have
a large infrastructure change. Um all right, incident response. Uh, this is your disaster recovery playbook. It's your business continuity playbook. You know, what would happen if I got hit by a bus? I have a lot of Nexus knowledge in my head. I probably have more than I should. Uh, you know, what's going to happen? Who else can push live? Who else can access uh the different environments?
So, you need to make sure that all of that is planned and documented out. Um, when there's an incident, take care of the incident. But then at that point, log everything that's happening, you know, as you go. And after the incident, have a review. Um, you know, don't wait a month to say, "Okay, let's finally talk about that thing that happened." Try and figure out what happened.
You don't need to blame anybody, but you definitely don't want to do, "Let's just do better next time." Find what the issue was and have actionable items to resolve that. um you want to test these uh disaster recovery and business continuity plans uh you know at least once a year and it's just a tabletop game of uh nerdy computer D and D you know this person uh
is unavailable now or the system went down or we just got ransomware let's open up the playbook let's follow the process uh so that everybody knows what they're supposed to do when they're Uh if your real incident is your first practice, then you're already behind. Uh when it comes to this, um there are a ton of controls, so many controls uh that you have to cover. You
don't have to cover the ones uh that you're not responsible for. I don't have to prove that AWS data centers have locks on the doors. They've already done that. They're going to share their SOC 2 type 2 report with me. And there's a shared responsibility matrix that I can give my auditors that says we're hosted here. that's everything they're doing and I inherit all of their passing
controls automatically. Um, so you can avoid uh duplicating that evidence gathering and leverage those provider certifications. Um, every uh security uh certification is going to require a pentest in one form or another. Um, I kind of treat this like I clean if I were ever to hire someone to come clean my house, I'm going to clean it first to get all the bad stuff done uh and
cleaned up and then let that person uh clean the nitty-gritty. Um, there's lots of tools you can use to run uh pentest on your own platform. Fix the obvious stuff. Uh, obviously you want to fix anything critical and high severity. Um we use uh za proxy um uh because our main endpoints are the our web app our APIs and our infrastructure and offflow. So that is the
main things that we tested. Um the pentest came through our uh a company dual joint with our uh consulting uh that helped us with the security. um you define the rules of engagement um and you know just test it and once it's done you need to fix any issues it brings up specifically critical and high issues because if you don't fix them and have it tested again
to show that they were fixed the auditors will treat it as if it wasn't done and you will get dinged hard for that. um asset inventory tracking, my laptop, my work phone. You want to maintain an inventory of that. It can be a spreadsheet to begin with. It can be endpoint management as long as it's documented somewhere. As long as it's not in somebody's head, track ownership,
where is it, you know, what's the life cycle of it? Um tag it. Uh that's pretty straightforward. Uh now, let's talk about more of the Drupal side of things. uh with the few minutes we have left, Drupal is great and if you treat it like a framework uh and not just a CMS out of the box, um you can you know prepare and build in tools early
uh for your compliance. Auditors are always going to be asking can you prove it uh configuration management is your evidence trail. It's it's if it's not in git then you know it didn't happen. Least privileged means you can hand over roles and permission matrix without you know wincing that oh man look at all these people that have admin access with huge permissions uh against them. Um there
are a lot of modules that help support uh security and compliance in different ways. We've seen some of them. Uh here are a few. Again everybody will have these slides. um inputs and forms have modules that help with uh you know captas antibbots autobands uh your configuration and supply chain you know we get an email anytime there's a update that's required uh our composer audit and mpm
audit in the CI conflict split if you have APIs and you're exposing JSON API to people do they need to get the entire entity object do they just need a few fields on that object JSON API extras can let you limit the scope of the data that's going out. You know, we've talked about the DevOps code quality items um previously. Uh when you're building custom modules, build
them with compliance in mind. Make sure you're logging security relevant actions anytime they happen in your code. Use those entity access handlers, not just random permission checks everywhere. Emit structured events and not like subscribed events, but contextual logs. um not just a a watchdog string message that said this happened. Add context to it. Um I said that one twice. Drupal has a new access policy API um
which is really great for uh defining access rules declaratively instead of scattering logics across different hooks. Um it makes them a little bit easier to audit. Uh and we already talked about that one. Uh multi-tenant, we talked about that one already. the evidence you can get from Drupal. Um, roles and permissions, audit logs. Uh, we're going to skip ahead with my last uh, thing here to the
blueprint. So, how does all of this fit together? Um, a complianceaware architecture starts with your process, your policies, your procedures, your documentation, your onboarding, your offboarding checklists, your access reviews. All of that is the process definition. that is your governance and personal sec personnel security and the change justification your infrastructure that everything's built on top of uh is has your account separation your SCPs all everything's encrypted
that's your config management your lease privileged your code flow your PRs your branch protections your security scans your automated tests there's all of your change management controls your vulnerability management controls your input and output validation your pipeline the CI/CD with gates your tagged releases, your deployment logs. There's your separation of duties and your audit trail of how did this actually get to production. Your obser uh observability
dashboards, alerts, and log aggregation lets you know before the observation period that your stuff is working, during the observation period that your stuff is working. Uh and after that period that it continues to work and the evidence you provide is everything above that you've just built, the policies, the screenshots, the configurations. There's a lot of consultants that can tap straight into AWS and Azour in Google and
just query and see that oh yeah all of their users in AM have multiffactor authentication. I don't even need a screenshot of that because we queried it directly. Um so you know find a consultant that has those types of integrations. The goal is continuous readiness. Every commit and deploy uh is you know an access change is logged in one form or another. That's what separates compliance as
a project from compliance as engineering. You built all of this through the talk. This blueprint is the system working together. Um, so thank you everybody. You know, uh, I made this initial meme when there was 153 slides. I'm now up to 187. Biggest takeaway is document everything continuously as you go. Integrate this into your daily uh, work life. You need to start early. This is a long
journey to get through and the more habits you build now uh will just make uh delivering compliance as easy as dumping logs later and taking a few screenshots. Know what's in scope. You don't have to be compliant on everything absolutely everywhere. Just the critical things. Uh treat it like engineering. Automate as much as possible. Put it in version control. Make sure you're reviewing it continuously and always
be improving it. And of course, continue to document everything you do continuously. Thank you very much. Uh we are out of time, but if you want to stay and ask some questions, I will keep this going for a couple more minutes. >> Logging Drush. uh automatic Drush. So the question is do we have any solutions for logging Drush? Ideally you're not arbitrarily running Drush commands in your
critical production environments. Any commands that you run you know for deployment are going to be captured via that log. Um with that said there are solutions out there. Uh I know that tail scale has one built into SSH monitoring um that it will keep a log and monitor who's connecting to your servers and every command that was run and output. tail scale. >> Tail scale is what
we use, but there's other systems that do the same. >> Is there anything regarding blue green deployment strategies something? >> So the question is there is there anything uh regarding blue green deployments? Um that's great. I've always had issues with blue green deployments with Drupal especially when you have to run you know a Drush update script. you it's hard to have you know two databases running at
the same time in that environment but um that blue green deployment would be part of that pipeline of um uh QA and UAT and securely rolling out and monitoring you know if you have the ability to do that that's where all of that fits in um I don't we don't use it because we usually have a lot of database updates and it's practically impossible to do that
and we have a massive database Yeah, >> exactly. Uh any thoughts on saving keys within like uh I know you didn't use GitLab but uh it's like CI/CD variables or is like vault or some kind of secrets management software like an absolute must. >> Yeah. So you know so the question is is there any thoughts on saving uh secrets in CI/CD uh environment variables? Um, every CI/CD
has, you know, a built-in secret environment man, uh, management section for that purpose. Um, that's the right thing to do at a minimum. Um, that's perfectly acceptable. Your policy says there's no secrets in code. You store them in the, uh, CI/CD configuration. Uh, if you have the ability to, you know, have it automatically rotate and sync with other places, that's the nec uh, next best step. But
there's nothing wrong with using those secrets. uh in that fashion. >> How do you take secrets back? Is it custom? >> How do you what? >> Back up secrets. >> How do you back up secrets? Um we rely uh we back up our secrets in one password. Um so that's how you know they're let's say we're using CI/CD in this case and we get a secret from
uh you know the server I don't know whatever it is it goes in the CI/CD. um that's where it lives primarily, but then it's also backed up uh in one password. So if for some reason that configuration gets deleted or changed or whatever, we have it stored in a separate place as well. Um that's the manual way to do it. Um some of these tools like AWS
um uh key manager and the AWS backup vault, it has the ability to back up your secrets for you. >> Did you have a question? Do you have any recommendations for SSO tools that work globally including China? >> Um, so the question is do I have any globally including China? Um I I don't uh have expertise there um beyond if there's an SSO tool that works globally
including China then you know and it it uses that same uh what is it uh OIDC protocol then you should be fine using one of the modules that support that in Drupal uh but I don't have expertise there beyond definitely use SSO if all right I'm going to hit the button