# Bommarito Consulting, LLC — Full Content > Expert advisory at the intersection of AI, governance, information security, privacy, and computational modeling. Website: https://bommaritollc.com Contact: inquiry@bommaritollc.com --- ## About Bommarito Consulting, LLC is a professional consulting firm offering expert advisory at the intersection of AI, governance, information security, privacy, and computational modeling. Founded in 2011, the firm operates at the intersection of technology, governance, and strategy — helping organizations make informed decisions in complex, fast-moving domains. The team brings credentials including CPA, CIPP/US, CIPP/E, and Certified AI Auditor, backed by a combined research portfolio of nearly 50 academic publications with 3,800+ citations. ### Focus Areas - Artificial Intelligence: Strategy, implementation, auditing, and governance - Board & Executive Advisory: Strategic guidance, board roles, and C-suite advisory - Governance & Risk: AI governance, privacy compliance, risk management - Expert Testimony: Testimony, speaking, and education - Computational Modeling: Environmental and economic modeling, simulation, and analysis - Information Security: Security advisory and compliance frameworks ### Timeline - **2011:** Bommarito Consulting founded, providing technology and strategy advisory services. - **2014:** Published Supreme Court prediction research using machine learning — one of the first applications of AI to legal outcomes forecasting. - **2014–2018:** Advised on $10B+ in capital events, working with Fortune 50 companies, major law firms, and financial institutions. - **2013:** Co-founded LexPredict, building AI-powered legal technology tools for contract analysis and litigation analytics. - **2018:** LexPredict acquired — successful exit from legal technology venture. - **2022:** Team earns Certified AI Auditor credential in the first cohort, before ChatGPT launches. - **2024:** GPT bar exam research published in Philosophical Transactions of the Royal Society A, demonstrating AI passing the Uniform Bar Exam. First 'Fairly Trained' certified LLM launched with 132M+ copyright-clean documents. - **2025+:** Expanding advisory practice with continued investment in open legal and economic AI research. --- ## Team ### Michael Bommarito **Role:** Managing Member Michael is a serial entrepreneur, researcher, educator, and advisor with 25 years of experience at the intersection of artificial intelligence, governance, and computational modeling. He has published nearly 50 articles and books in venues including Science, Philosophical Transactions of the Royal Society A, and Cambridge University Press, with nearly 3,800 academic citations. His pioneering work on GPT passing the bar exam has been widely covered in mainstream and legal media. As an entrepreneur, Michael has co-founded, scaled, and exited multiple technology companies, and has helped raise and deploy over $10B in private capital. He has worked with many of the largest corporations, law firms, and financial institutions. Michael continues to lead active research in open legal and economic AI, and oversees development of large-scale data infrastructure for responsible LLM training — including assembling over 132 million copyright-clean documents across government, legal, and financial domains. **Highlights:** - Nearly 50 publications with 3,800+ academic citations - Published in Science, Phil. Trans. Royal Society A, Cambridge University Press, PLOS ONE, Journal of Statistical Physics - Pioneered systematic AI evaluation on professional exams (bar exam, CPA exam) - Built the KL3M Data Project — 132M+ copyright-clean documents, 1.35 trillion tokens for LLM training - Created LexNLP, an open source NLP library for legal text - Developed SCOTUS prediction models achieving 70.2% accuracy across 240,000+ justice votes - Maintainer of FOLIO — 18,000+ standardized legal concepts across 10 languages - Author: This Is Server Country, The Math Inside the Machine, AI for Law and Finance - Fastcase 50 legal innovator (2019) - NSF Fellow (IDEAS-IGERT) **By the Numbers:** - Publications: ~50 - Academic Citations: 3,800+ - Projects: 53 - Years Experience: 25+ **Focus areas:** Artificial Intelligence, Governance & Strategy, Computational Modeling, Information Security **Open to:** Research collaborations, Speaking & educational opportunities, Expert testimony & consulting, Advisory & board positions **Education:** [object Object]; [object Object]; [object Object] **Experience:** President, ALEA Institute; CEO, 273 Ventures; CEO, LexPredict (acquired 2018); CodeX Affiliate, Stanford Center for Legal Informatics; Head of Research, Law Lab, Chicago-Kent College of Law; Adjunct Professor, Michigan State University College of Law; Lecturer, Center for the Study of Complex Systems, University of Michigan **Awards & Recognition:** - Fastcase 50 (2019) - NSF Fellowship, IDEAS-IGERT (2009–2010) - Google Summer of Code (2007, 2008, 2009) **Links:** [Website](https://michaelbommarito.com), [GitHub](https://github.com/mjbommar), [LinkedIn](https://linkedin.com/in/bommarito), [Google Scholar](https://scholar.google.com/citations?user=gIDHq_oAAAAJ&hl=en) ### Jillian Bommarito **Role:** Principal **Credentials:** CPA, CIPP/US, CIPP/E, Certified AI Auditor Jillian is a risk and governance expert with 15 years of experience across startup, corporate, and consulting environments, specializing in financial compliance, privacy, and AI auditing. She brings a rare combination of CPA, privacy, and AI auditing credentials, and has held leadership roles at multiple ventures at the intersection of technology and compliance. She oversaw the development of the first 'Fairly Trained' certified large language model, which was subsequently open sourced through a nonprofit. She is adept at distilling complex technical concepts for boards, regulators, and executive teams. Jillian has been featured in Wired, BBC World Service, Law360, AICPA Journal of Accountancy, and the Canadian Bar Association National Magazine. She co-authored research on AI CPA capabilities and contributed to empirical surveys published in the Virginia Tax Review. She currently serves as Executive Director and Head of Governance at the ALEA Institute, where she oversees responsible AI development practices, and as Chief Risk Officer at 273 Ventures. She previously founded and led licens.io, a technology diligence and valuation firm for software and data assets. **Highlights:** - First cohort Certified AI Auditor (2022, before ChatGPT launched) - Oversaw training of the world's first 'Fairly Trained' LLM - Featured in Wired, BBC World Service, Law360, AICPA Journal of Accountancy - ALM Female Founders in LegalTech honoree - Women in AI Governance (WiAIG) Regional Chapter Chair - Co-authored GPT as Knowledge Worker CPA evaluation research - Co-authored empirical survey published in the Virginia Tax Review **By the Numbers:** - Years Experience: 15+ - Professional Credentials: 4 - Media Features: 10+ **Focus areas:** AI Governance & Auditing, Privacy (US & EU), Risk Management, Financial Compliance **Open to:** Speaking & educational opportunities, Advisory & corporate board positions, Limited consulting engagements, Collaborating on open research **Education:** [object Object]; [object Object] **Experience:** Executive Director & Head of Governance, ALEA Institute; Chief Risk Officer, 273 Ventures; CEO, licens.io; CFO, LexPredict (acquired 2018) **Awards & Recognition:** - ALM Female Founders in LegalTech - Women in AI Governance Regional Chapter Chair **Speaking:** Mozilla + Eleuther Dataset Convening on Open Datasets for AI (2024), Deep Dive into AI for Legal, Doon Insights (2024), 16th Annual IT Law Seminar, State Bar of Michigan (2023), Fin(Legal)Tech Conference, Chicago-Kent College of Law (2017), LawNext with Bob Ambrogi (2024), AI for Smart People (2024), The Law of Tech (2024), The Geek in Review (2024) **Media:** Wired, BBC World Service, Law360, AICPA Journal of Accountancy, Canadian Bar Association National Magazine, ALM Legal Tech News, Benzinga **Links:** [Website](https://jillianbommarito.com), [LinkedIn](https://linkedin.com/in/jbommarito-cpa-cipp-us), [X](https://x.com/jillbommar) --- ## Services ### Advisory & Board Strategic guidance through board positions and advisory roles for organizations navigating AI, governance, and technology strategy. We serve on boards of directors and advisory boards for organizations that need experienced guidance at the intersection of technology, governance, and strategy. Our advisory engagements range from early-stage startups defining their AI strategy to established organizations navigating regulatory change and digital transformation. **Offerings:** - **Board of Directors Positions:** Active board participation with governance, technology, and strategy oversight. - **Advisory Board Membership:** Strategic guidance on AI, data strategy, and open source decisions. - **AI Governance Committee:** Specialized guidance on responsible AI adoption, risk frameworks, and policy. - **Strategic Technology Guidance:** CTO-level advisory on technology architecture, build vs. buy, and roadmap. - **Risk and Compliance Oversight:** Financial, privacy, and AI risk review from credentialed professionals. **Relevant credentials:** CPA, CIPP/US, CIPP/E, Certified AI Auditor **Related services:** technology-strategy, ai-governance ### Technology Strategy & Security AI strategy, technology due diligence, M&A technical assessment, and information security advisory for organizations making high-stakes technology decisions. We help organizations make high-stakes technology decisions — from AI adoption strategy and enterprise security posture to M&A technical due diligence and compliance readiness. Our team combines hands-on experience building and scaling technology companies with deep expertise in information security, financial engineering, and AI systems. Every engagement is grounded in real operational experience across the full company lifecycle, from founding through acquisition. We assess technology risk, architecture, and strategy from the perspective of people who have built, governed, and exited the systems under review. **Offerings:** - **AI Strategy & Readiness:** Enterprise AI adoption roadmaps, use case prioritization, vendor evaluation, and deployment planning. - **Technology Due Diligence:** Technical assessment of platforms, architecture, and engineering teams for investment and strategic decisions. - **M&A Technical Assessment:** Pre-acquisition technology, security, and IP evaluation for private equity, venture capital, and corporate buyers. - **Information Security Advisory:** Security posture review, framework alignment (ISO 27001, SOC 2), and remediation planning. - **Security & Compliance Readiness:** Enterprise sales readiness, security questionnaire programs, and compliance framework implementation. - **Enterprise AI Readiness:** Organizational readiness assessment for AI adoption, covering data governance, vendor risk, and change management. **Relevant credentials:** M.S.E. Financial Engineering, CPA, CIPP/US **Related services:** advisory, ai-governance, privacy-compliance ### Expert Testimony & Speaking Expert testimony, speaking engagements, and educational sessions on AI, governance, and technology. Our team provides expert testimony and speaking engagements grounded in years of hands-on experience building, researching, and advising on AI systems and technology platforms. We bring credibility through published research, real-world implementation experience, and recognized credentials in AI auditing, privacy, and financial compliance. **Offerings:** - **Expert Witness Testimony:** Testimony on AI capabilities, technology standards, and industry practices. - **Conference Keynotes and Panels:** Engaging presentations on AI, governance, and technology trends. - **Corporate Education Sessions:** Board and executive briefings on AI adoption and risk. - **Regulatory and Legislative Briefings:** Expert input for policy development and regulatory proceedings. - **Academic Guest Lectures:** Research-backed instruction for law, business, and technology programs. - **Podcast and Media Appearances:** Accessible expert commentary on AI, technology, and governance. **Related services:** advisory, technology-strategy ### AI Governance & Auditing Responsible AI frameworks, governance committees, and independent AI auditing backed by first-cohort certification. We help organizations implement responsible AI practices through governance frameworks, independent auditing, and committee advisory. Our work is grounded in hands-on experience overseeing the development of the first 'Fairly Trained' certified LLM. Our Certified AI Auditor credential — earned in the first cohort in 2022, before ChatGPT launched — combined with deep technical expertise in building AI systems, enables us to evaluate AI risk from both governance and engineering perspectives. **Offerings:** - **AI Governance Framework Development:** Design and implement governance structures for responsible AI adoption. - **Independent AI Auditing:** Third-party evaluation of AI systems for bias, risk, and compliance. - **AI Ethics Committee Advisory:** Guidance on forming and operating effective AI oversight committees. - **Responsible AI Policy Development:** Organizational policies for AI development, procurement, and deployment. - **AI Risk Assessment:** Systematic evaluation of AI-related risks across the organization. - **Fairly Trained Certification Support:** Guidance on copyright-clean training data practices and certification. **Relevant credentials:** Certified AI Auditor, CPA, CIPP/US **Related services:** advisory, technology-strategy, privacy-compliance ### Privacy & Data Protection Comprehensive privacy programs for US and EU requirements, backed by CIPP/US and CIPP/E certifications. We develop and implement privacy programs that address the full spectrum of US and European data protection requirements. Our team holds both CIPP/US and CIPP/E certifications, providing authoritative guidance on GDPR, CCPA/CPRA, and emerging state and international privacy regulations. Our privacy practice is informed by direct experience building data-intensive technology products, including technology diligence platforms and large-scale document collections for responsible AI training. We understand privacy not just as a compliance exercise, but as a technical and operational discipline. **Offerings:** - **Privacy Program Development:** End-to-end design and implementation of organizational privacy programs. - **GDPR Compliance:** Assessment, gap analysis, and remediation for EU data protection requirements. - **CCPA/CPRA Compliance:** California privacy law compliance, including consumer rights workflows. - **Data Protection Impact Assessments:** Systematic assessment of processing activities and privacy risks. - **Data Licensing and Rights Management:** Frameworks for data licensing, consent, and intellectual property. - **Privacy by Design Implementation:** Embedding privacy controls into product and system architecture. **Relevant credentials:** CIPP/US, CIPP/E, CPA **Related services:** ai-governance, technology-strategy ### Computational Modeling Environmental and economic modeling, simulation, and quantitative analysis for complex decision-making. We build and apply computational models to help organizations understand complex systems — from environmental risk and climate impact to economic forecasting and market dynamics. Our work combines deep quantitative expertise with practical domain knowledge to produce actionable insights. Our modeling practice integrates expertise in financial engineering, applied mathematics, and large-scale data analysis. Whether the challenge is simulating regulatory impact, modeling supply chain risk, or quantifying environmental exposure, we deliver rigorous, defensible analysis. **Offerings:** - **Environmental Modeling:** Climate risk, emissions modeling, and environmental impact simulation. - **Economic & Financial Modeling:** Market dynamics, regulatory impact analysis, and quantitative forecasting. - **Agent-Based Simulation:** Complex systems modeling using agent-based and Monte Carlo simulation techniques. - **Risk Quantification:** Probabilistic modeling of operational, financial, and environmental risk. - **Scenario Analysis:** Multi-scenario planning and stress testing for strategic decisions. - **Data Pipeline Development:** End-to-end data infrastructure for model inputs, calibration, and reporting. **Related services:** technology-strategy, advisory ### Fractional Executives On-demand C-suite leadership for startups, scaling companies, and organizations navigating pivots — without the overhead of a full-time hire. Not every organization needs — or can afford — a full-time executive team. We provide fractional C-suite leadership that brings senior strategic guidance on the schedule and budget that fits your stage. Our fractional executives have built, scaled, and exited technology companies, hold credentials across finance, privacy, AI governance, and information security, and bring the perspective of operators who have sat in the seats they advise. **Offerings:** - **Fractional CTO:** Technology strategy, architecture decisions, build vs. buy, engineering team leadership, and technical due diligence. - **Fractional CFO:** Financial strategy, capital planning, financial controls, audit readiness, and transaction support. - **Fractional CRO:** Revenue strategy, go-to-market execution, sales process optimization, and growth planning. - **Fractional CIO:** IT strategy, digital transformation, vendor management, and enterprise systems governance. - **Fractional Chief AI Officer:** AI strategy, responsible AI governance, vendor evaluation, and organizational AI readiness. - **Fractional CISO:** Security posture, compliance readiness, incident response planning, and risk management. **Relevant credentials:** CPA, CIPP/US, CIPP/E, Certified AI Auditor, M.S.E. Financial Engineering **Related services:** fractional-cfo-cro, fractional-cto-cio-caio, advisory ### Fractional CFO & CRO Fractional CFO and CRO leadership for organizations that need senior financial strategy, revenue operations, and compliance oversight without a full-time hire. Growing companies face a gap between bookkeeping and the strategic financial leadership that boards, investors, and acquirers expect. Our fractional CFO and CRO services bridge that gap with experienced leadership grounded in real operational, audit, and transaction experience. Our team brings CPA credentials, M&A experience on both sides of the table, and a track record across startup finance, corporate governance, and revenue operations — from 409A compliance and audit committee support to go-to-market strategy and capital planning. **Offerings:** - **Financial Strategy & Planning:** Budgeting, forecasting, cash flow management, and financial modeling for growth and capital decisions. - **Audit & Compliance Readiness:** Financial controls, SOX readiness, audit committee support, and regulatory compliance for finance teams. - **Transaction & Capital Support:** Due diligence, 409A valuations, capital raise preparation, and M&A financial advisory. - **Revenue Strategy & Operations:** Go-to-market planning, sales process optimization, pricing strategy, and revenue forecasting. - **Board & Investor Reporting:** Financial reporting packages, KPI dashboards, and board-ready financial communications. - **Risk & Financial Governance:** Enterprise risk frameworks, financial policy development, and internal controls design. **Relevant credentials:** CPA, M.S.E. Financial Engineering **Related services:** fractional-executives, fractional-cto-cio-caio, advisory ### Fractional CTO, CIO & Chief AI Officer Fractional technology and AI executive leadership — CTO, CIO, and Chief AI Officer — for organizations making high-stakes technology and AI decisions. Technology leadership requires more than technical skill — it requires judgment informed by real experience building, governing, and exiting technology companies. Our fractional CTO, CIO, and Chief AI Officer services provide that judgment on a schedule that matches your needs. We bring deep expertise in AI systems, information security, and computational infrastructure — backed by nearly 50 publications, 3,800+ academic citations, and hands-on experience across the full company lifecycle from founding through acquisition. **Offerings:** - **Fractional CTO:** Technology strategy, architecture review, engineering leadership, build vs. buy decisions, and technical due diligence. - **Fractional CIO:** IT governance, digital transformation roadmaps, vendor strategy, enterprise systems, and infrastructure planning. - **Fractional Chief AI Officer:** AI strategy development, responsible AI governance, model evaluation, vendor assessment, and organizational AI readiness. - **AI Governance & Risk:** AI risk frameworks, bias evaluation, compliance with EU AI Act and NIST AI RMF, and responsible AI policy. - **Security & Compliance:** Security posture review, ISO 27001 and SOC 2 readiness, and information security governance. - **Technology Due Diligence:** Technical assessment of platforms, architecture, and AI systems for investment, M&A, and strategic decisions. **Relevant credentials:** M.S.E. Financial Engineering, Certified AI Auditor, CIPP/US, CIPP/E **Related services:** fractional-executives, fractional-cfo-cro, ai-governance, technology-strategy --- ## Insights ### EU AI Code of Practice Draft: What Boards Should Know **Date:** 2024-11-14 | **Category:** AI Governance An analysis of the EU AI Office's draft General-Purpose AI Code of Practice, highlighting systemic risk oversight, documentation requirements, incident response, and board-level compliance considerations. The EU AI Office has released a draft General-Purpose AI Code of Practice. While not legally binding, it is expected to be influential in shaping the AI landscape in the EU. Boards of General-Purpose AI providers should be aware of the implications for their organizations. The Code of Practice covers transparency, copyright rules, taxonomy of systemic risks, safety and security frameworks, risk assessment, technical risk mitigation, and governance risk mitigation. Importantly, the AI Act applies to AI models provided for free — including open source models — as well as those that are sold. Sub-Measure 15.2 explicitly calls on boards to establish oversight of systemic risks from general-purpose AI models, including through the creation of dedicated risk committees. The Code establishes the need for board-level responsibility for allocating adequate resources for overseeing systemic risks, including ensuring executives have sufficient budgets and the right expertise. Organizations are required to document their adherence to the Code and all applicable provisions of the AI Act. This includes technical documentation of AI models, criteria for classification, security and safety framework documentation, and evidence collected during risk assessments. Risk assessment is a key focus, with four Measures relating specifically to how providers should assess systemic risks — continuously, from before training through post-deployment. The Code also requires documented incident response plans and establishes whistleblower protections under the EU Whistleblower Directive. Boards should keep an eye on the Code's development. While it may change following the public comment period, many of the Practices align with broader best practices for AI governance. If your board does not have sufficient expertise to address AI governance, consider bringing in an additional board member or engaging in board-level training specific to AI risks. ### AI Lifecycle and the Board's Role **Date:** 2024-10-11 | **Category:** AI Governance A primer for board directors on the AI lifecycle — data collection, training, and deployment — and the strategic considerations boards must understand for effective AI oversight. Boards are expected to make critical strategic decisions about their companies' approaches to AI. It's tough to do that without a basic understanding of the AI development lifecycle and the drivers of AI systems. This guide helps board members understand the high-level steps of training a model and how to think about data, software, and hardware when developing strategic plans. The AI model lifecycle generally progresses through three steps: data collection, training and fine-tuning, and deployment and integration. The quality of training data has a significant impact on model performance — garbage in, garbage out still applies. Board considerations for data collection include data governance, data provenance and IP rights, data protection frameworks, and competitive advantage from proprietary data. Most organizations won't train a foundation model from scratch but will fine-tune an existing model. Board considerations for training include resource allocation and scalability, alignment with business objectives, environmental impact on ESG objectives, ethical guidelines including bias mitigation, intellectual property strategy, and talent acquisition. Once deployed, models deliver value in real-world situations. Deployment requires ensuring security and performance both independently and jointly with other systems. Board considerations include stakeholder management, performance monitoring with defined KPIs, and incident management with board-level reporting mechanisms. Three key assets drive the AI lifecycle: data, software, and hardware. Each is subject to changing economics and trends. The value of proprietary, domain-specific datasets has been demonstrated by organizations like Bloomberg, which trained models on its own financial data and outperformed general models on domain-specific tasks. One of the best ways to address the changing nature of AI is to craft a strategy that is flexible and adaptable. Ensure that products and vendors allow for data portability, so your organization can switch to better solutions as they emerge. By understanding the interplay between data, software, and hardware, boards are better able to guide strategy, ensure responsible adoption, and make informed decisions about resource allocation. ### Risk Management for AI: A Board Director's Guide **Date:** 2024-10-09 | **Category:** AI Governance A comprehensive guide for board directors on leading AI risk management through six key elements: establishing context, risk assessment, risk treatment, recording, communication, and continuous monitoring. AI risk management starts with the board. Board directors play a critical role in setting the firm's risk appetite, establishing context and objectives, and ensuring that AI initiatives align with overall strategy. Well-established risk management frameworks like ISO 31000, COSO ERM, and the NIST Cybersecurity Framework can be adapted for AI-specific risks. The risk management process begins with establishing context and objectives — understanding the current regulatory environment and the specific uses of AI within the organization. This includes data flow mapping, a critical step in risk identification that helps boards understand jurisdictional data flows and regulatory requirements. Risk assessment encompasses identification, analysis, and evaluation. AI introduces unique risk categories including bias, opacity, data provenance challenges, and rapidly evolving regulatory requirements. Boards must ensure assessment processes are comprehensive and regularly updated. Risk treatment options include avoidance (choosing not to deploy AI in high-risk contexts), mitigation (implementing technical guardrails and human oversight), transfer (through insurance or contractual allocation), and acceptance (where risks fall within the organization's defined tolerance). Each approach has implications that boards should understand. Recording and reporting should be carefully documented to support transparency and accountability. By creating and preserving records, organizations can demonstrate compliance with regulatory obligations and industry standards. Despite mitigation strategies, risks may still be realized — having a well-crafted response plan minimizes impact. Effective risk management is continuous, not a one-time exercise. Given the speed at which AI and related legal obligations evolve, boards must regularly evaluate their AI risks and opportunities. This involves assessing whether the existing program still meets organizational needs, identifying new risks, finding areas for improvement, and implementing changes that strengthen the program's effectiveness. ### AI Oversight: 5 Key Sources of Board Requirements **Date:** 2024-10-10 | **Category:** AI Governance A framework identifying the five key sources of AI governance requirements for boards — legal mandates, risk frameworks, insurance, internal policies, and customer preferences. While procurement decisions are made below the board level, the decision about if and how an organization will use AI falls within the strategic and governance oversight of the board. Requirements and constraints generally come from five different sources. First, legal and regulatory requirements — the most obvious source. Companies operating globally face complex compliance challenges. Second, risk management frameworks — sometimes incorporated into law (e.g., NIST publications), but organizations may also adopt them independently. Third, insurance requirements — many professional liability providers now require policyholders to disclose how their organizations use AI. Fourth, internal policies and economics — driven by external forces or by board and management preferences. Economic considerations may constrain how organizations can realistically procure and deploy AI. Fifth, customer and partner preferences — customers may request specific jurisdictional processing, limiting which models or products an organization can utilize. Common requirements across these sources include data governance and security (encryption at rest and in transit, data processing rules, retention and deletion policies), access control and monitoring (authentication, authorization, audit trails), and operational resilience (business continuity and disaster recovery planning). Human resources and third-party management requirements include personnel vetting, training, and third-party vendor management — especially important for AI solutions that integrate open source software and multiple service providers. Risk management requirements include insurance practices and intellectual property rights policies. By understanding these common requirements, board directors can provide more effective oversight of AI initiatives. These requirements should be viewed not just as compliance hurdles but as areas where boards can add value through strategic guidance and risk management. ### How Data Provenance Drives Machine Learning Risk and Value **Date:** 2022-03-30 | **Category:** AI Governance An exploration of data provenance — the origin and history of data — and why it is a board-level concern for AI risk management, legal compliance, and responsible governance. Provenance is just knowing where something came from. In art, provenance answers two questions: Who created the work? And does the current possessor have the right to transfer? In the context of data, provenance answers: What individuals or organizations are described in the data? And does the current possessor have the right to transfer or use? Answering these questions is critical for three reasons. First, knowing your data: how much value can you create from data if you don't understand or trust it? When you acquire data second-hand or third-hand, trust becomes increasingly important. Second, contractual obligations: contracts may explicitly prohibit re-using or re-distributing data, and downstream use creates breach of contract risks. Third, regulatory compliance: federal and state laws may require specific documentation with respect to data. Neglecting data protection regulations can have severe impacts on an organization's machine learning models, operations, and financial wellbeing. If consent is the basis for processing data, there should be a means by which data can be tracked to consent. Copyright considerations also play a significant role in data usage rights. If data has been obtained from a third party, the issue of provenance is even more important. Organizations should document the lineage of third-party data and any limitations on that data. Technology is both enabling the exponential growth of data — which complicates provenance — and offering potential solutions. MLOps platforms like MLflow allow organizations to create and version datasets, including their provenance and lineage, and use these datasets to train versioned machine learning models. Data provenance is not just a technical concern but a fundamental aspect of responsible AI governance. For board members, understanding and overseeing data provenance practices is crucial for ensuring the reliability, legality, and ethical use of AI and machine learning models. In the world of AI and data, knowing where your data comes from is just as important as knowing where you're going with it. ### Pioneering Responsible Data Science: A Framework for Ethical Innovation **Date:** 2021-10-18 | **Category:** AI Governance Introducing an open-source Responsible Data Science Policy Framework designed to help organizations address ethical AI governance through a modular, adaptable approach. Less than half of the world's largest organizations have governance procedures for AI ethics in place. This gap is particularly concerning given increasing scrutiny from regulators and investors worldwide. Regulators in the US and EU are drafting new rules for data science practices, while investors managing over $35 trillion in assets are incorporating data science-related ESG considerations into funding criteria. Recognizing this need, we developed a Responsible Data Science Policy Framework through Licensio. The framework addresses several key challenges: risk management (identifying and mitigating risks associated with data science and AI), legal compliance (staying ahead of emerging regulations), ethical considerations (ensuring responsible use of data and AI), and stakeholder trust (building trust with customers, investors, and the public). What sets this framework apart is its modular and adaptable design. It consists of a parent procedure for triaging specific use cases, prescriptive sub-procedures for low-friction compliance, and adjudicative sub-procedures for centralized decision-making. This structure allows organizations to start with a basic committee-based approach and evolve towards more specialized processes over time. There are five key reasons organizations should care. First, strategic oversight and competitive advantage: a robust governance framework sets you apart. Second, risk mitigation: proactively addressing ethical and legal concerns prevents costly mistakes. Third, innovation enablement: clear guardrails actually enable faster, more confident innovation. Fourth, stakeholder confidence: demonstrating responsible practices builds trust with customers, investors, and regulators. Fifth, future-proofing: as regulations evolve, a flexible framework makes adaptation easier. We open sourced this framework to make it accessible to all organizations committed to responsible data practices. Whether you're a startup just beginning to leverage AI or a large corporation with established data practices, this framework offers a roadmap to more ethical, efficient, and valuable data science operations. ### The KL3M Data Project: Building Copyright-Clean Training Data at Scale **Date:** 2024-09-01 | **Category:** Open Source Inside the KL3M Data Project — assembling 132M+ copyright-clean documents to enable responsible LLM training and earning the first 'Fairly Trained' certification. The KL3M Data Project represents one of the most ambitious efforts to assemble copyright-clean training data for large language models. With over 132 million documents sourced from public domain, government, and openly licensed materials, the project demonstrates that it is possible to train capable AI systems without relying on copyrighted content. The project emerged from a simple observation: as organizations and regulators increasingly scrutinize the provenance of AI training data, there is a growing need for training datasets with clear intellectual property rights. The legal risks associated with training on copyrighted material — from litigation to regulatory action — create a compelling case for copyright-clean alternatives. Building KL3M required solving several technical challenges. Document sourcing at scale, quality filtering, deduplication, and format normalization all required significant engineering effort. The project also required careful legal analysis to ensure that each data source met our copyright-clean standards. The result was not just a dataset but a demonstration that responsible AI training is feasible. When an LLM trained on KL3M data received the first 'Fairly Trained' certification, it validated the approach and established a new benchmark for training data governance. For organizations navigating AI adoption, the KL3M experience offers practical lessons. Data provenance is becoming a first-order concern — not just for legal risk, but for customer trust, regulatory compliance, and competitive positioning. The firms and institutions that invest in responsible data practices now will be better positioned as governance standards tighten. The KL3M Data Project is maintained through the ALEA Institute and continues to grow. It stands as an example of how open source and nonprofit collaboration can address systemic challenges in AI development. ### Since Our Last Episode: The Evolution of Bommarito Consulting **Date:** 2019-01-02 | **Category:** Firm Update A retrospective on the firm's journey from early-stage consulting through the LexPredict era, and the pivot toward AI governance, privacy, and institutional advisory. When Bommarito Consulting was founded in 2011, the landscape of AI and legal technology looked fundamentally different. The firm began as a general technology consulting practice, but quickly found its niche at the intersection of artificial intelligence, legal systems, and governance. The intervening years brought a series of pivotal developments. The founding of LexPredict in 2013 represented a bet on the commercial viability of AI-powered legal technology — a bet that paid off with a successful acquisition in 2018. The experience of building, scaling, and exiting a legal technology company informed everything that followed. By 2019, the firm's focus had sharpened considerably. Rather than broad-spectrum technology consulting, we found our highest-value work at the intersection of AI governance, privacy compliance, and institutional advisory. Our clients increasingly needed help not just adopting technology, but governing it responsibly. The shift wasn't just strategic — it was credentialed. The team's CPA background combined with emerging privacy certifications (CIPP/US, CIPP/E) and eventually one of the first Certified AI Auditor credentials positioned the firm to serve a growing market. Continued research output, including the landmark GPT bar exam study, ensured our advisory work remained grounded in frontier knowledge. Looking ahead from that vantage point, the firm was poised for what would become an era of unprecedented demand for AI governance expertise. The launch of ChatGPT in late 2022 would transform the market — but we had been preparing for years. ### Built to Sell: Lessons from the LexPredict Journey **Date:** 2018-06-15 | **Category:** Entrepreneurship Key takeaways from building, scaling, and successfully exiting LexPredict — an AI-powered legal technology company acquired in 2018. Building a company with the intention of creating lasting value — whether that means an exit, sustained growth, or something in between — requires a fundamentally different mindset than building a lifestyle business. The LexPredict experience taught us several lessons that continue to inform our advisory practice. First, domain expertise matters enormously. LexPredict succeeded in part because the founding team brought genuine expertise in both legal technology and artificial intelligence. We weren't technologists learning about law or lawyers learning about technology — we were researchers and practitioners who had spent years at the intersection. Second, timing and market dynamics are critical. LexPredict launched at a moment when the legal industry was beginning to take AI seriously but before the market was saturated with competitors. The company's early investment in contract analysis and litigation analytics positioned it well for acquisition. Third, intellectual property and research credibility create durable competitive advantages. LexPredict's published research, open source contributions, and academic affiliations weren't just marketing — they established genuine authority that customers and acquirers valued. Fourth, the importance of governance and compliance cannot be overstated, even in a startup context. Having a CPA as CFO, maintaining clean financials, and implementing proper data handling practices made the due diligence process dramatically smoother. These lessons directly inform our advisory work today. When we sit on boards or advise executives, we bring the perspective of operators who have navigated the full lifecycle — from founding through exit — not just consultants who theorize about it. ### Predicting the Supreme Court: AI Meets Legal Outcomes **Date:** 2017-03-10 | **Category:** Research How our machine learning research achieved breakthrough results in predicting Supreme Court decisions, and what it means for the future of legal AI. In 2014, we published research demonstrating that machine learning models could predict the outcomes of U.S. Supreme Court cases with meaningful accuracy. The work, which appeared in PLOS ONE and was subsequently covered by major media outlets, represented one of the earliest applications of AI to legal outcomes forecasting. The approach was deceptively simple: by training models on historical case features — the issue area, the lower court decision, the ideological composition of the Court, oral argument characteristics — we could predict case outcomes at rates significantly above baseline. The model's performance on out-of-sample data provided strong evidence that the patterns were real, not artifacts of overfitting. What made the research significant wasn't just the accuracy numbers — it was the demonstration that legal outcomes, often perceived as the product of pure reasoned judgment, contained statistical regularities that machines could detect. This finding had implications for legal strategy, judicial analytics, and the broader question of how law operates in practice. The Supreme Court prediction work also illustrated a theme that would define much of our subsequent research and advisory practice: the intersection of AI capabilities with professional domains. The same questions we asked about judicial prediction — How accurate can AI be? Where does it fail? What are the ethical implications? — later resurfaced in our work on AI bar exam performance and CPA evaluation. For Bommarito Consulting, this research established a foundation of credibility that continues to differentiate our advisory practice. When we advise organizations on AI capabilities and limitations, we do so as researchers who have published in the field, not merely as consultants who read about it. ### Course Material for Complex Systems 530 — Computer Modeling for Complex Systems **Date:** 2015-01-27 | **Category:** Research Open course material for Complex Systems 530 at the University of Michigan, covering agent-based, Monte Carlo, and network modeling in Python. Complex Systems 530 — Computer Modeling for Complex Systems at the University of Michigan Center for the Study of Complex Systems. In the spirit of open science, all course material is available online at Github: https://github.com/mjbommar/cscs-530-w2015. The course explores why and how we model the world around us, from an interdisciplinary perspective using Python. The goal is to help students understand how to frame and formulate models, understand and select appropriate modeling methodologies (including agent/individual-based, Monte Carlo, and systems/structural models), implement and analyze models using Python, and communicate model methodology and results. The course IPython notebooks are available via NBViewer. Try the following to get started: Monte Carlo and deforestation, basic grids and the Schelling model of segregation, and basic networks and disease outbreak models. If you'd like to discuss similar training for your data science or modeling team, please don't hesitate to reach out. ### Predicting the Behavior of the Supreme Court of the United States: A General Approach **Date:** 2014-07-27 | **Category:** Research Introducing our Supreme Court prediction project using extremely randomized trees to forecast over sixty years of decisions and individual justice votes. One of the more exciting and public projects we've been working on lately has finally come to light — our Supreme Court prediction project with Dan Katz and Josh Blackman. This project is exactly what you'd expect — a framework for predicting the Supreme Court, though meant to span the Court's entire history, unlike previous projects. While we'll be releasing more information on the methodology, results, and next steps in the coming days, stop by the project page, read the paper, or review the abstract below. Building upon developments in theoretical and applied machine learning, as well as the efforts of various scholars including Guimera and Sales-Pardo (2011), Ruger et al. (2004), and Martin et al. (2004), we construct a model designed to predict the voting behavior of the Supreme Court of the United States. Using the extremely randomized tree method first proposed in Geurts, et al. (2006), a method similar to the random forest approach developed in Breiman (2001), as well as novel feature engineering, we predict more than sixty years of decisions by the Supreme Court of the United States (1953-2013). Using only data available prior to the date of decision, our model correctly identifies 69.7% of the Court's overall affirm/reverse decisions and correctly forecasts 70.9% of the votes of individual justices across 7,700 cases and more than 68,000 justice votes. Our performance is consistent with the general level of prediction offered by prior scholars. However, our model is distinctive as it is the first robust, generalized, and fully predictive model of Supreme Court voting behavior offered to date. Our model predicts six decades of behavior of thirty Justices appointed by thirteen Presidents. With a more sound methodological foundation, our results represent a major advance for the science of quantitative legal prediction and portend a range of other potential applications. Katz, Daniel Martin and Bommarito, Michael James and Blackman, Josh, Predicting the Behavior of the Supreme Court of the United States: A General Approach (July 21, 2014). Available at SSRN: http://ssrn.com/abstract=2463244 ### Advanced Approximate Sentence Matching in Python **Date:** 2014-06-30 | **Category:** Archive Advanced techniques for approximate sentence matching in Python using Jaccard similarity on token, stem, and noun lemma sets after stopword removal. In our last post, we went over a range of options to perform approximate sentence matching in Python, an important task for many natural language processing and machine learning tasks. To begin, we defined terms like tokens (a word, number, or other discrete unit of text), stems (words that have had their inflected pieces removed based on simple rules), lemmas (words that have had their inflected pieces removed based on complex databases), and stopwords (low-information, repetitive, grammatical, or auxiliary words that are removed from a corpus before performing approximate matching). We used these concepts to match sentences via exact case-insensitive token matching after stopword removal, exact case-insensitive stem matching after stopword removal, and exact case-insensitive lemma matching after stopword removal. These methods all work by transforming or removing elements from input sequences, then comparing the output sequences for exact matches. If any element of the sequence differs or their orders change, the match fails. In order to locate more true positives, we need to relax our definition of equivalence. The next most logical way to do this is to swap our exact sequence matching with a set similarity measure. One of the most common set similarity measures is the Jaccard similarity index, which is based on the simple set operations union and intersection. We are comparing two sentences: A and B. We represent each sentence as a set of tokens, stems, or lemmas, and then we compare the two sets. The larger their overlap, the higher the degree of similarity, ranging from 0% to 100%. If we set a threshold percentage or ratio, then we have a matching criterion to use. Let's say we define sentences to be equivalent if 50% or more of their tokens are equivalent. When we parse sentences to remove stopwords, we end up with sets like {young, cat, hungry} and {cat, very, hungry}. The Jaccard index numerator counts shared items (2: cat, hungry) and the denominator counts total unique items (4: young, cat, very, hungry). Therefore, our Jaccard similarity index is 2/4 = 50%. In Approach #4, we apply case-insensitive token set similarity after stopword removal. Instead of looking for an exact match between the sequence of tokens, we calculate the Jaccard similarity of the token sets. This allows us to survive small insertions or order changes in the sequence. In Approach #5, we replace tokens with stems, which allows us to equate words with a common meaning. By switching from tokens to stems, we match additional sentences that differ only in word inflection. In Approach #6, we replace stems with lemmas and also ignore all lemmas that do not correspond to nouns. The motivation is that we care about the "things" in the sentence, not the "actions" — using only noun lemma sets allows us to achieve more positive matches while controlling false positives. ### Fuzzy Match Sentences in Python **Date:** 2014-06-12 | **Category:** Archive A tutorial on fuzzy matching sentences in Python using NLTK, covering tokenization, stopword removal, stemming, and lemmatization approaches. Let's imagine you have a sentence of interest. You'd like to find all occurrences of this sentence within a corpus of text. How would you go about this? The most obvious answer is to look for exact matches. But what if capitalization, punctuation, or white-spacing varied in the slightest? Consider the sentence: "In the eighteenth century it was often convenient to regard man as a clockwork automaton." Small variations in capitalization, whitespace, or punctuation would cause exact matching to fail, even though the substance of the sentence is identical. We need to learn to fuzzy match sentences, not exact match sentences. To perform fuzzy sentence matching, we need to move from matching exact strings to more flexible, natural-language-motivated definitions of equality. Examples include: exact case-insensitive token matching after stopword removal, exact case-insensitive stem matching after stopword removal, exact case-insensitive lemma matching after stopword removal, and partial set similarity approaches. Our first improvement is case-insensitive token matching after stopword removal. This means ignoring case, treating the sentence as a sequence of tokens, and ignoring stopwords (high-frequency, low-content words like "is", "or", "the"). After processing our example sentence, we get: ['eighteenth', 'century', 'often', 'convenient', 'regard', 'man', 'clockwork', 'automaton']. This blurs whitespace, punctuation, case, and low-content words. The next approach is stemming — the process of "uninflecting" or "reducing" words to their stem. In English, common examples include "cats" → "cat" and "printing" → "print". Stemming is typically implemented using preset rules that may not handle irregular words. The Porter and Snowball stemmers, for example, fail on irregular nouns like "children" and "women". The third approach uses lemmatization, which takes a more complex but comprehensive approach. The WordNet lemmatizer handles irregular forms correctly: "children" → "child", "women" → "woman". One important source of complexity is that lemmatization relies on part-of-speech tagging — "printing" as a noun is not inflected, but "printing" as a verb should reduce to "print". Once we've executed lemmatization, we can handle cases like the Greek plural "automata" being correctly matched to "automaton". At this point, we've learned a lot about tokenizing, stopwording, stemming, and lemmatizing, but we've only matched about half of the sentences that we would characterize as similar. The next post covers partial matches using Jaccard set similarity. ### Isotonic Regressions in scikit-learn **Date:** 2014-06-08 | **Category:** Archive A practical example and discussion of isotonic regression in Python's scikit-learn, including contributions to improve the IsotonicRegression class. Isotonic regression is a great tool to keep in your repertoire; it's like weighted least-squares with a monotonicity constraint. Imagine that the true relationship between x and y is characterized piece-wise by a sharp decrease in y at low values of x, followed by a gradual decrease in y for larger x. There is also heteroskedasticity, with greater errors at low values of x. The key is that we believe y should be strictly non-increasing in x, i.e., monotonic. Now imagine we want to produce a smoothed or denoised version of this relationship for visualization or to regularize data for further modeling. Common smoothing methods like polynomial splines, LO(W)ESS, or non-linear least squares won't obey our monotonicity assumption without additional work. Isotonic regression, on the other hand, is explicitly designed for this purpose. In practical applications, we are probably trying to use observed values of y to predict some further z. We want to use past experience about x and y to help us better predict z. Isotonic regression handles this naturally while preserving the monotonic constraint. Under the hood, this is handled in Python by scikit-learn's IsotonicRegression class. I recently pushed a few enhancements to the IsotonicRegression class: PR 3157 automatically determines whether y is increasing or decreasing in x based on the Spearman correlation coefficient; PR 3199 handles out-of-domain x values gracefully instead of throwing ValueError exceptions; and PR 3250 provides efficiency improvements by storing the interpolating function at fit time. You can follow along with the Python code in the IPython notebook: http://nbviewer.ipython.org/urls/gist.githubusercontent.com/mjbommar/6b355ecfeb60051c799c/raw/isotonic-regressions-in-scikit-learn.ipynb. ### Is the Tax Code the Longest Title? **Date:** 2013-08-19 | **Category:** Research An analysis of whether the Internal Revenue Code (Title 26) is actually the longest Title in the U.S. Code, using data from our paper on measuring legal complexity. Last week, I shared that Dan Katz and I had finally published a draft of our paper, Measuring the Complexity of the Law: The U.S. Code. Since then, we've received great feedback and a number of questions. The most common question, even among legal professionals, is exactly what you'd guess — is the Tax Code (i.e., Title 26, I.R.C.) the longest Title? The answer, in our opinion, is also what you might guess — it depends. But first, let's look at a few measures of Titles in the Code. These plots are all based on data from our Github repository and the source to reproduce them can be found in the accompanying R and ggplot2 gist. What do we notice? Title 26 is not the longest or biggest by any measure. It doesn't have the most words (Title 42), the most elements/sections (again, Title 42), or even the most words per section (Title 23). So what can we say? Are Titles the right unit of measure? Titles are the first cut of the hierarchical categorization of the U.S. Code. It is generally accepted that they do not always represent a cohesive body of law; for example, Title 42 — Public Health and Welfare, is an amalgamation of topics as diverse as commercial space transportation, farm housing, and healthcare. However, with Acts, they are the most commonly discussed group. Is any division of the Code atomic? If you've read any statutory text, you are familiar with references or citations that incorporate definitions, rules, or other language. If Title 26 and Title 42 are heavily interdependent through reference, does it make sense to compare them? We believe the only proper way to do this is by incorporating measures of the network structure of the Code. Hopefully, this discussion has piqued your interest in measuring legal complexity and raised your awareness around some common pitfalls. If so, please give our paper a read and let us know what you think! ### Measuring the Complexity of the Law: The U.S. Code **Date:** 2013-08-13 | **Category:** Research Releasing our empirical framework for measuring legal complexity, applied to the U.S. Code, with full replication source and data on GitHub. Four years ago, Dan Katz and I began working on a project to measure the complexity of the law. Its genesis was, in every sense, an accident; in order to properly identify citations to the IRC in our empirical review of U.S. Tax Court decisions, we had to deal with the informal, non-Blue Book citation standard used by Tax Court judges. To do this, we scraped and parsed out all possible subsection citations to the IRC and identified these citations in written opinions. When we were done, we had a directory hierarchy that corresponded to titles, parts, chapters, and sections of the U.S. Code. Immersed in our NSF Fellowships at the University of Michigan Center for the Study of Complex Systems, we asked the contextually obvious question — could we measure the complexity of this system? Since then, this question has taken many forms: text and XML, LaTeX and Word, Physica A and CELS. The full paper, intended for broader consumption, has gone through hundreds of hours of major and minor revisions. We kept sitting on the paper because it was never perfect, because there was always another refinement, because we always had another idea. Last Thursday, this paper finally saw its first "official" release. It's still a draft, and we still haven't decided which publication venue to pursue, but we felt it was important to contribute to the recent discussion around open Federal legislative material. As part of the paper, we've released all necessary replication source (Python) and data on Github. You can view the repository here: mjbommar/us-code-complexity. Einstein's razor, a corollary of Ockham's razor, is often paraphrased as follows: make everything as simple as possible, but not simpler. This rule of thumb describes the challenge that designers of a legal system face — to craft simple laws that produce desired ends, but not to pursue simplicity so far as to undermine those ends. Complexity, simplicity's inverse, taxes cognition and increases the likelihood of suboptimal decisions. While many scholars have offered descriptive accounts or theoretical models of legal complexity, empirical research to date has been limited to simple measures of size, such as the number of pages in a bill. In this paper, we address this need by developing a proposed empirical framework for measuring relative legal complexity. This framework is based on "knowledge acquisition," an approach at the intersection of psychology and computer science, which can take into account the structure, language, and interdependence of law. ### Summary of Community Detection Algorithms in igraph 0.6 **Date:** 2012-06-17 | **Category:** Research A reference guide to the community detection algorithms available in igraph 0.6, including their runtime complexity, and support for directed and weighted edges. Based on Launchpad traffic and mailing list responses, Gabor and Tamas will soon be releasing igraph 0.6. In celebration, I'll be publishing a number of helpful lists and tables I've put together to organize information about igraph. In this post, we'll cover the community detection algorithms (i.e., clustering, partitioning, segmenting) available in 0.6 and their characteristics, such as their worst-case runtime performance and whether they support directed or weighted edges. Much of the information below is gleaned from the igraph C documentation, source algorithm publications, and three years of tracking the 0.6 trunk. Optimal Modularity — Modularity is more of a framework than just a simple function over a graph. The amount of work based on this idea, implicitly or explicitly, is staggering given the short six years since. That said, modularity is just a framework, and, like all frameworks, has its shortcomings. The performance of modularity maximization in practical contexts behaves poorly in real networks, and, depending on |V| and |E|, we might not see "smaller" clusters. Details: New to 0.6, undirected only, unweighted only, handles multiple components, runtime B^(|V|^2). Edge-Betweenness — Community structure in social and biological networks (Girvan and Newman). Supports directed edges, weighted edges, handles multiple components, runtime |V||E|^2. Leading Eigenvector — Finding community structure in networks using the eigenvectors of matrices (Newman, 2006). Undirected only, unweighted only, handles multiple components, runtime c|V|^2 + |E|. Fast-Greedy — Finding community structure in very large networks (Clauset, Newman, Moore). Undirected only, supports weighted edges, handles multiple components, runtime |V||E| log |V|. Multi-Level — Fast unfolding of communities in large networks (Blondel, Guillaume, Lambiotte, Lefebvre). New to 0.6, undirected only, supports weighted edges, handles multiple components, "linear" runtime when |V| ≈ |E|. Walktrap — Computing communities in large networks using random walks (Pons, Latapy). Undirected only, supports weighted edges, does not handle multiple components, runtime |E||V|^2. Label Propagation — Near linear time algorithm to detect community structures in large-scale networks (Raghavan, Albert, Kumara, 2007). New to 0.6, undirected only, supports weighted edges, does not handle multiple components, runtime |V| + |E|. InfoMAP — The map equation (Rosvall, Axelsson, Bergstrom). New to 0.6, supports directed edges, supports weighted edges and nodes, does not handle multiple components, estimated runtime |V|(|V| + |E|). ### Connecting R to an Oracle Database with RJDBC **Date:** 2012-11-22 | **Category:** Archive A step-by-step guide to connecting R to an Oracle database using RJDBC, a cross-platform JDBC-based approach. In many circumstances, you might want to connect R directly to a database to store and retrieve data. If the source database is an Oracle database, you have a number of options: ROracle, RODBC, and RJDBC. Using ROracle should theoretically provide you with the best performing client, as this library is a wrapper around the Oracle OCI driver. The OCI driver, however, is platform-specific and requires you to install Oracle database client software. Using RODBC for Oracle is like using an ODBC connection for any database; so long as your platform provides an ODBC manager and drivers, you are OK. On Linux, this means unixODBC, and on Windows, this means the Oracle Data Access Components package. What if you don't want to write code that is either platform-specific or requires relatively complex, platform-specific installation steps? In this case, you should consider using RJDBC. I'll assume that you have a JRE/JDK installed and know the path to your JAVA_HOME. The first step is to obtain the Oracle JDBC drivers, e.g., the 11gR2 release drivers. You can pick the lowest compatible Java version you'd like to support; I'm using ojdbc6.jar, which should support Java 6+. Next, make sure you know how to connect to your source database. You'll need the following information for your database listener: hostname or IP, port (e.g., 1521), service name or SID (e.g., ORCL), username, and password. This information will allow us to construct the DSN, which will look something like this: jdbc:oracle:thin:@//hostname:port/service_name_or_sid. ### Saving Memory in Redis and Python with struct.pack **Date:** 2012-01-15 | **Category:** Archive Using Python's struct.pack to convert numbers from string to binary representation in Redis, achieving significant memory savings. In redis, every object is either a binary-safe string or collection thereof. Even if you're storing numbers from your client library or using the INCR/DECR counter within redis, the actual values are stored in memory as strings. In some cases, this may be fine; however, if you are storing many numbers, whole or decimal, this may add up. If you're using Python to store and retrieve data from redis, there's an easy way to save loads of memory — the built-in struct library. struct handles converting Python types to binary representations of underlying C types. In the case of numbers, this means that we can quickly and painlessly convert that sequence of characters into the underlying binary representation, which is quite often smaller. The simple pattern packs non-negative whole numbers into the smallest binary format: 'B' for numbers under 2^8 (unsigned char), 'H' for numbers under 2^16 (unsigned short), 'I' for numbers under 2^32 (unsigned int), and 'L' for larger values (unsigned long). In a simple test of 1M integers uniformly distributed between 1 and 10^5, this pattern resulted in a 45% reduction in memory usage. While there's no hard-and-fast rule, I've had at least 25% reductions in almost every real-world application of this technique. YMMV, but given the ease of this hack and the cost of memory, it's definitely worth a shot! If you have a better handle on the range or structure of data and aren't afraid of math, you should also look at rolling your own encoding with GETRANGE or GETBIT. More to come on this in a future post. ### 21st Century Legal Informatics: Part 1, Introduction **Date:** 2011-11-13 | **Category:** Research An introduction to three paradigms of legal informatics — 20th century computers-as-libraries, 22nd century computers-as-lawyers, and the practical 21st century middle ground. Dan and I have written and spoken on legal informatics many times. Inevitably these conversations come to the same list of informatics examples from legal search/retrieval and decision making. These examples fall into two categories. The first sits firmly in the 20th century, while the second belongs in the 22nd century. I'll support my argument below and conclude with a lead into what I'd like to call 21st century law. 20th Century Legal Informatics — Computers as Libraries. Ask a typical lawyer how informatics affects their practice, and they might mention that salary infographic their friend emailed them. Data, modeling, statistics, and visualization might be seen as cute toys, but not real tools. Ask a typical lawyer how search affects them, however, and they'll readily list five-figure-per-seat services like Lexis, West, CCH, or RIA. My opinion is that search is the only informatics tool that fits into the current legal paradigm — the library model. Law is a field of humans interpreting words, words live on documents, and documents live in libraries. Legal training focuses on reading and interpreting words and documents. For a new tool to be accepted by lawyers, it must complement this library model. From this standpoint, it's easy to see why search has succeeded — it facilitates the traditional library model. 22nd Century Legal Informatics — Computers as Lawyers. This category is best understood through IBM Watson and the International Association for Artificial Intelligence and Law (IAAIL). Watson embodies the hopes of 22nd century legal informatics, in which computers build and interpret models to make legal decisions. The IAAIL and its members have been presenting data models, search methods, expert systems, and judicial reasoning for more than 30 years. However, as a participant and former member, I will readily admit that the IAAIL has mostly failed to introduce these ideas into the mainstream of legal practice. 21st Century Legal Informatics — Computers and Lawyers. What can we do while our robotic overlords are still incubating? The way forward rests on four principles: Balance (neither humans nor computers alone provide optimal outcomes), Measure but not too much (measure wherever possible but avoid promoting measurement when it isn't the solution), Change but not too much (focus on cases where accurate models can be built, such as finance), and Aim high (don't let expectations based on the library model set your bar for success). ### Building a Better Legal Search Engine, Part 1: Searching the U.S. Code **Date:** 2011-04-10 | **Category:** Research An introduction to indexing and searching the U.S. Code using Apache Lucene, structured public domain data, and open source software. The first part in a blog series leading up to a keynote on Law and Computation at the University of Houston Law Center focuses on indexing and searching the U.S. Code with structured, public domain data and open source software. Before diving into the technical aspects, some background on the U.S. Code. After a bill is passed and becomes Public Law, it is published in the Statutes at Large — the authoritative, chronological compilation of all enacted law. However, the Statutes at Large is sorted by date of enactment, not by concept; it contains laws that may affect multiple legal concepts, reference other laws, and amend or repeal other laws. This makes exhaustive search impractical. The U.S. Code, produced by the Office of the Law Revision Counsel, solves this by organizing the law by concept (hierarchically), combining laws that reference one another, removing expired or repealed laws, and providing convenient citations. The LRC distributes copies of the Code in XHTML, which we use to build our index. To build a legal search engine, the Code is arguably the best place to start. While there are other important sources like the Code of Federal Regulations or the Federal Reporter, the Code is as close to capital-L Law as it gets. We use the Apache Lucene library to build an index of the Code from the 2009 and 2010 LRC snapshots. Lucene is a high-performance, full-featured text search engine library written entirely in Java. The process took a little over two minutes on a laptop. The buildCodeIndex.java program extracts XHTML files, tokenizes documents into sections, and passes them through Lucene's analyzer (which stems tokens and strips stopwords). Once the index is built, a simple single-term search interface (searchCodeIndex.java) allows querying. A search for "swap" across the entire Code returns top results including sections on swap dealer registration, swap recordkeeping, swap execution facilities, and swap data repositories — all highly relevant to post-Dodd-Frank compliance. ### Tracking the Frequency of Twitter Hashtags with R **Date:** 2011-02-21 | **Category:** Archive An R script for downloading and plotting the frequency of Twitter hashtags over time using ggplot2. I've posted three examples of Twitter hashtags datasets in the last week: one on China, one on Iran, and one on Algeria. In order to build these datasets, I needed to obtain older tweets; this is slightly more difficult than simply filtering the streaming feed for your hashtag of choice. The original code I wrote for this task is in Python and is well-parallelized, but the code isn't commented and looks more complicated than it is due to parallelization choices. As part of my recent exercise to replace Python with R for entire tasks, I decided to rewrite this code using R. The code is pretty simple, well-commented, and consists of two functions — loadTag and downloadTag. There is one significant issue with the code: at the moment, neither rjson nor RJSONIO seem to support Unicode data in JSON responses. Furthermore, when character vectors of unknown encoding are written to file with a function like write.table, they produce output that cannot be reliably read back into R. As a result, the code does not retain the text of a tweet — only the id, date, and username. Once you've downloaded some data, producing frequency figures is only two lines away with ggplot2: load the tweets with loadTag, then plot with geom_bar using 5-minute binwidths. ### Plotting 3D Graphs with Python, igraph, and Cairo **Date:** 2011-02-21 | **Category:** Archive A demonstration of how to generate 3D network animations using Python, igraph, and Cairo, applied to a Twitter user graph. Out of all the visuals I've produced, I think the "coolest" is the three-dimensional U.S. Supreme Court citation network 1080p movie I produced with Dan Katz. 3D networks, especially dynamic ones, really invoke the "wow" factor. Movies are especially important in dynamic cases, since without the animation, you lose all context and understanding of what led the network to that point. The 3D part may seem unnecessary at first, but an extra dimension can go a long way to unwinding the hairball that most somewhat-dense networks form. How does it work? Using a Twitter dataset, along with Python, igraph, and Cairo to demonstrate how 3D network animations can be generated. In this example, we animate the rotation of a static snapshot of a user graph. The program works as follows: First, build a network based on @username mentions in the tweet dataset. Then, calculate a three-dimensional Kamada-Kawai layout of (x,y,z) tuples. While sweeping the polar and azimuthal angles, calculate a two-dimensional projection of the currently rotated layout, draw the vertices and edges in increasing z-order, and output the surface. If we wanted to dynamically add and remove vertices or edges from this network, changes would be needed to re-calculate and interpolate layouts across network snapshots. The graphmovie library for Python (http://code.google.com/p/graphmovie/) is designed to do exactly this. The Python code is available as a gist at Github (https://gist.github.com/837396). ### Historical Data Mining the Supreme Court Headnotes **Date:** 2011-05-04 | **Category:** Research A demonstration of historical data mining techniques using editorially-assigned legal headnotes from U.S. Supreme Court cases, examining co-occurrence and citation networks to trace the evolution of legal concepts like criminal suspect rights. Two weeks ago, I posted a pair of very rough working papers. The second of these, Exploring Relationships between Legal Concepts in the United States Supreme Court, opens up a number of interesting "historical data mining" techniques. I thought I'd go over an example of this today to demonstrate the usefulness of the approach. First, let's go over the paper at a very high level. In order to facilitate legal research, legal publishers like LexisNexis and West tag cases from their available ontology of legal concepts. These headnotes are very helpful, especially when dealing with "multi-dimensional" cases where the Court may address (or argue why it should not address) more than one legal question. Examples of Lexis's headnotes include: Constitutional Law > Bill of Rights > Fundamental Freedoms > Freedom of Speech > Forums; Administrative Law > Agency Adjudication > Alternative Dispute Resolution; Workers' Compensation & SSDI > Third Party Actions > Third Party Liability; Criminal Law & Procedure > Interrogation. The first thing you should notice is that these headnotes are hierarchical. There are top level categories, like Constitutional Law, Administrative Law, and Criminal Law & Procedure, as well as lower level categories, like Interrogation, Third Party Liability, and Forums. These concepts occur at different levels, and the figure below conveys the overall structure of this concept hierarchy. There are 42 separate top level head notes, 603 second level headnotes, and many more below. Since 42 headnotes is probably too little categorization to say much about a Supreme Court case, let's focus on second level headnotes. These give us 603 separate, editorially-assigned tags for cases. Each case can have 0 or more of these tags. So how might we procede? First, let's look at how often these headnotes co-occur within a case. Co-occurrence can imply a number of things. For example, it might mean that the facts of a case led two separate legal issues to come into question. Alternatively, it might mean that a Justice used analogical reasoning to "import" precedent from one legal concept to another. Regardless of the specific reason, the frequencies of these co-occurrences broadly indicate the level of interaction between legal concepts. To visualize the macro-level structure of these relationships, we can examine the weighted network layout below. Alternatively, we could look at citation instead of co-occurrence. In this case, we are interested in cases with one headnote that cite a case with another headnote. Like co-occurrence, these citations may imply multiple relationships. For example, just as above, these could indicate instances of "concept importation" where precedent from one domain is applied to another. The weighted network visual below displays the resulting citation relationships among headnotes. These visuals are interesting, but how could we use these ideas to ask specific legal questions? As a case study, let's try to trace the history of criminal suspect rights such as those discussed in Miranda v. Arizona. Let's say that we are particularly interested in how the Bill of Rights was used in discussion of criminal interrogation. The standard research approach might look like this: determine a set of terms or phrases that correspond to discussion of the Bill of Rights; determine a set of terms or phrases that correspond to discussion of criminal interrogation; combine these in a query to Lexis or West; reverse sort by date; examine each case to determine whether your query did a good job; adjust query and repeat until happy. With headnotes, however, we can simply count the number of cases where the Bill of Rights headnote and the Interrogation headnote co-occur. The first case that matches this logic is Hopt v. People of the Territory of Utah, 110 U.S. 574. "Elementary writers of authority concur in saying that, while from the very nature of such evidence it must be subjected to careful scrutiny and received with great caution, a deliberate, voluntary confession of guilt is among the most effectual proofs in the law, and constitutes the strongest evidence against the party making it that can be given of the facts stated in such confession. 1 Greenleaf Ev. § 215; 1 Archbold Cr. Pl. 125; 1 Phillips’ Ev. 533-34; Starkie Ev. 73." A time series representation of these co-occurrences in the figure below also gives us meaningful information about these issue over time. While the result may not seem surprising given how much attention has been paid to development of Miranda rights, this technique is equally useful in less-studied domains. If any of these ideas have seemed interesting, feel free to check out the paper on SSRN. There's plenty more information and visualization in the paper. Bommarito, Michael James, Exploring Relationships between Legal Concepts in the United States Supreme Court (November 5, 2009). Available at SSRN: http://ssrn.com/abstract=1814169 ### Building Legal Language Explorer: Interactivity and Drill-Down, noSQL and SQL **Date:** 2011-12-16 | **Category:** Research A technical deep-dive into the architecture of the Legal Language Explorer, a Google Ngrams-style viewer for the U.S. Supreme Court corpus, combining redis (noSQL) for fast time series queries with PostgreSQL for case-level drill-down. Dan and I recently released a new legal informatics project with a few colleagues. The project, which we've named the Legal Language Explorer, provides an interface similar to Google Ngrams Viewer for the U.S. Supreme Court. Unlike Google's viewer, however, the Legal Language Explorer also allows users to drill-down into case-level information for each n-gram. The technical architecture that allowed us to robustly provide both of these worlds wasn't simple, but I felt that the experience and result are worth sharing. Before I get into the technical aspects, I wanted to mention one especially important feature of the project – the low cost of failure in interaction. By engineering a system where the downside of bad choices is small relative to the payoffs, people are much more willing to experiment. Sure, that search for "gorilla" was a dud, but it only took 50ms. Before disappointment can even register, you've already gone on to "mickey mouse" or "Kant." This may seem like an unimportant point, but when search spaces are large, exploring them requires asymptotically small costs. For Legal Language Explorer, this low cost is a direct consequence of combining two technologies – SQL and noSQL. (Aside: While SQL and noSQL are often viewed as antithetical, I think this is a product of a confused conversation. Our real goal should be to store or process data efficiently. noSQL is a response to the frustrations of attacking problems like graph traversal, document store, or tall-and-skinny numeric with traditional, row-store SQL. The world isn't so black-and-white that we can generalize well, however. For example, SQL databases like Vertica can compete with noSQL databases like kdb for financial data, and SQL databases that support WITH RECURSIVE can compete with neo4j on shallow traversals.) The front page of Legal Language Explorer displays a time series plot of usage for "interstate commerce," "railroad," and "deed" between 1791 and 2005. This plot contains 645 (year, n-gram) values and represents a total of 103,285 occurrences. If your experience is like most people's, the plot has loaded before the page text has even rendered. Is it just a cached PNG? Is the data hard-coded into the page? The answer lies in redis, a small, neatly written key-value store that many refer to as a noSQL database (to me, redis is more of an in-memory data structure server than a database). Regardless of what you call it, redis excels at providing blazingly fast read and write access to common structures. It does this by keeping all data live in memory, organized for efficient access. Here are some basic facts and figures on our redis instance: redis Architecture: Single 64-bit instance; Memory: 12GB running (4GB dump); Keyspace: 103,281,603 keys, up to 214 fields per key; Average Keyspace Hit Rate: 77.7%. Despite the size of these numbers, the HGETALL query that drives the Legal Language Explorer plot returns in under 50ms for every query we've tested, even under load from 100 concurrent requests. Could you deliver a solution like that in a traditional RDBMS? If you find an n-gram that you want to learn more about, Legal Language Explorer allows you to drill down and view the list of all cases that n-gram occurs in. This table provides the full citation, year, and title for every case. Since there are often many parties with long names, storing this in redis would be expensive! What do we do? For drill-down, we rely on a traditional, row-store database; in this case, postgres. The recipe is fairly straightforward. Populate a (somewhat) normalized database with a case table, and an occurrence table, where occurrences are pairs of (ngram, case_id). Throw in an index on the n-gram, and you're done. The final product looks something like this: SELECT COUNT(*) FROM occurrences: 298,471,830; Indices: 10GB; Total Tablespace: 25GB. When we serve up data to the user, we SELECT on occurrences, constraining by n-gram, and JOIN the matching cases. (Would normalizing the n-grams help us get to 3NF? Yes, but it didn't help us conserve CPU while competing with redis.) The end result is that most queries are returned in under 10 seconds. This is great news for everyone who wanted a list of Supreme Court opinions referencing Elvis Presley. Search away! Legal Language Explorer allows users to explore n-grams in the Supreme Court corpus at the macroscopic and microscopic level. This dual mandate is met through a marriage of SQL and noSQL, and we feel that the whole experience is much greater than the sum of the parts. Over the next few months, we hope to raise enough money to add new data and features to the site, so please stay tuned! (If you're really paying attention to the geeky stuff, you may have noticed that I said nothing about the hardware we're running on. I didn't want to draw attention to this, since this post is about software, not hardware. However, if you have to know, the site is running on an EC2 m2.2xlarge spot instance for about $10/day. Not bad!) ### Statistics on the Length and Linguistic Complexity of Bills **Date:** 2012-02-13 | **Category:** Research An introduction to an automated daily analysis of 112th Congress bills, providing statistics on section count, word count, unique words, and Flesch-Kincaid reading level for every bill. Where would you go to find out what the longest bill of the 112th Congress was by number of sections (H. R. 1473)? How about by number of unique words (H.R. 3671)? What about by Flesch-Kincaid reading level (S. 475)? Head on over to this table of bills, updated daily for the 112th Congress, which contains the following fields: Bill Name, Publish Date, Bill Title, Stage, Section Count, Sentence Count, Word (Token) Count, Unique Words (Tokens), Unique Stem Count, Avg. Word Length, Avg. Sentence Length, and Reading Level (Flesch-Kincaid). I'll be adding more automated analysis and figures over the next few weeks, but for now, here's a morsel to get your gears turning. ### Building an AWS CloudSearch Domain for the Supreme Court **Date:** 2012-04-15 | **Category:** Research A step-by-step tutorial for building a fully searchable AWS CloudSearch domain using public domain U.S. Supreme Court decisions, covering data acquisition, domain configuration, indexing, and querying. It should be pretty clear by now that two things I'm very interested in are cloud computing and legal informatics. What better way to show it than to put together a simple AWS CloudSearch tutorial using Supreme Court decisions as the context? The steps below should take you through creating a fully functional search domain on AWS CloudSearch for Supreme Court decisions. Our first step is to acquire a public domain copy of Supreme Court decisions from Carl Malamud's resource.org. You can navigate to this directory and download US.tar.bz2, or just run something like: Once the download is done, extract the archive: We should now have a directory called US with 1.1GB and 62,839 files. Let's assume that you put this directory under something like /data/courts/US. The next step is easy – go follow my guide on setting up Cloud Search command line tools! I'll assume that you placed everything under /opt/aws/cloud-search-tools, just like in that post. OK, we should now have a dataset and the Cloud Search API at our fingertips. It's time to create a Cloud Search "domain" that we can populate with records. To do so, you can either follow the instructions on your AWS Management Console or run the following: This may take awhile to create; sometimes up to 15 minutes. Go grab a coffee or a beer and read your feed while you wait. You can check the status either through the Management Console in browser or with the following line: Once this step is complete, you should see an ACTIVE domain with 0 documents. We now need to reconfigure the access policies so that the domain allows us to submit search material and anyone to search: This policy change may take a few minutes to go into effect. Lastly, we need to tell the domain what we are indexing per document. OK, we're ready to go! At this point, we need to generate Search Data Format (SDF) files to populate the domain. There are two approaches we can take: (1) Write a parser to extract exactly the text content and metadata we want, or (2) Throw the pre-packaged cs-generate-sdf utility at our data and hope for the best. For brevity's sake, we'll pursue option 2. After some poking around, I've found that cs-generate-sdf is based on a common open-source content extraction library – Apache Tika. You might be familiar with Tika, as it's the guts behind Solr's ability to ingest unstructured data. So if you'd be happy naively ingesting the content in Solr, you'll probably be happy with the results that cs-generate-sdf produces. While we could build something more complex, let's stick to bash here: A few things to note: If you see error messages like "Request forbidden by administrative rules" or "403 Forbidden", your access policies have not taken effect or you provided the wrong IP for the document service. You should see lots of lines go by; two for every file that is being parsed. This step can be parallelized, but will almost certainly be disk-bound unless you are running on some kind of RAID or NAS setup that allows for concurrent reads. This could take awhile; about 45 minutes to generate and transmit on my i7 2600k/32GB RAM/SATA III SSD workstation. You should grab another coffee or beer and watch a show. Another caveat: even after you've transmitted all data up to the cloud, it will still take some time for the Cloud Search instance to churn through the data and complete indexing. Once the Cloud Search instance is fully built, it's time to figure out how to search. The best way to do this is, sadly, to read the developer documentation. However, if you want to skip all the boring part, just try running something like this: This search looks for an exact phrase match on "clear and present danger" and returns not only the document ID, but also the title property of the document. You should get back something like this: So, there it is! Your own fully searchable AWS Cloud Search domain for the Supreme Court. Not so bad after all, was it? ### “Google” for Subpoenaed Emails: AWS CloudSearch for eDiscovery **Date:** 2012-04-21 | **Category:** Archive A practical look at using AWS CloudSearch to build a scalable, on-demand search engine for subpoenaed email in eDiscovery engagements, eliminating the need for large capital expenditures on servers and storage. In the last post on AWS CloudSearch, I provided a tutorial on the creation of a simple CloudSearch domain for Supreme Court decisions. This walkthrough described the steps of creating a domain, configuring access policies and indexing, populating the index, and using the search API. We were left with a functioning case search database. From a technical perspective, one key difference between this example and many real-world applications is that we let the CloudSearch tools automatically decide what fields and content were available to search. While this worked well in the previous example, I want to provide a concrete example of a context in which custom services and development provide more value – eDiscovery. Imagine you’re a smaller law firm that specializes in HR disputes. As part of a time-sensitive non-solicitation claim filed by your client, you’ve subpoenaed email from fifteen employees at a client’s competitor. It’s Friday afternoon at 5PM, and you finally receive a hard drive with the emails. However, in an effort to overwhelm your small team, the competitor has produced the emails as individual RFC822 (.eml) files instead of a single Outlook PST file. You aren’t sure when you’ll be able to dig through the emails, and you already have a team of lawyers and legal assistants waiting. Combined with the right service provider (like Bommarito Consulting!), AWS CloudSearch is a perfect solution for this problem. Before CloudSearch, existing available on-site infrastructure constrained the provision of eDiscovery services. eDiscovery service providers had to make large capital expenditures on servers and storage to meet peak customer demand, and sometimes the capacity just wasn’t there when needed. This led to high fixed costs, which in turn forced providers to charge higher prices to customers. From the provider’s perspective, expensive server infrastructure was also underutilized for much of the year. CloudSearch makes these problems disappear. In our example above, building a “Google” for your subpoenaed emails can be done in just hours. The core components are an RFC822 parser to populate the search domain and a front-end user interface for searching and visualizing the results. If this service sounds valuable to your business, today or just potentially valuable in the future, don’t hesitate to reach out. ### eDiscovery Consulting in the Cloud: Searching an Outlook Mailbox and Attachments **Date:** 2012-05-19 | **Category:** Archive A real-world case study demonstrating how to make Outlook PST mailboxes and their attachments searchable using AWS CloudSearch, processing 1.3GB of Enron email data on a laptop in under an hour. You may have noticed that I keep talking about eDiscovery consulting and legal search in the cloud. I’ve covered searching the Supreme Court with new technologies in analytics and the cloud, making certain types of emails searchable on Amazon’s cloud, and even eDiscovery and the cloud at a high level. While these posts are all great teasers, you might be asking yourself, “What would it actually look like in practice?” I’d like to address that concern today by presenting a real-world case study on making Outlook PST mail and attachments discoverable. The software and process I will describe are fully implemented and tested. Nothing in this post is vaporware, and the entire process is scalable from a single mailbox up to an entire business. I’ll skip the usual disclaimers about privilege and process, and instead cut right to the point – what does the technology look like? In this example, I’ll be using 7 Outlook PST files shown below from the Enron email dataset. These mailboxes total 1.3GB and contain email text as well attachments in a variety of formats like PDF and Office. As such, they are a typical cross-section of corporate email. In order to convert this folder to a searchable database, we need to follow the following process: Examine each file in the folder and determine whether we understand the format. In this case, all 7 files are PST files, which we can process. For each of these mailboxes, we then build a list of all emails, storing information about who sent the message, who it was to, what they said, etc. For each email, we also identify any attachments. If we understand the attachment format, we also extract and store the textual content of the attachment and associate it with the related email conversation. Supported textual formats include MS Office 97–2010 Documents from Word, Excel, PowerPoint, etc.; Adobe PDF; OpenOffice Documents; Web pages; and plain or rich-text documents. Regardless of whether we can extract text from an attachment, we save the file to a folder on our computer where we can later examine it. We transmit all of this data to AWS CloudSearch, which handles indexing the textual content, as well as facets like from and to addresses. Once CloudSearch finishes building the index, we are ready to search! We have taken those 7 Outlook mailboxes and converted them into: a search interface for all textual email and attachment content, and a folder of attachment files embedded in all emails, such as images, audio, videos, or other non-textual content. In total, the total data set is processed on a laptop in under an hour, and search results and interfaces are available over the web almost immediately thereafter. If you’re interested in pricing or have any other questions, don’t hesitate to reach out. ### Now in Print: An Empirical Survey of the Population of U.S. Tax Court Written Decisions **Date:** 2011-04-15 | **Category:** Research Announcing the publication of an empirical analysis of U.S. Tax Court citation practices from 1990 to 2008 in the Virginia Tax Review, examining Internal Revenue Code citation frequency and the impact of tax legislation on court decisions. When someone brings up the empirical study of legal citation, most people think of the work Landes & Posner and Epstein & Martin. If you’re really cool, you might even think of Dan and me, who have spent quite awhile analyzing and visualizing Supreme Court citations like those in the 3D 1080p animation below: These studies, including our own Supreme Court papers, focus on just one particular type of citation, however – judicial citation. It’s important to remember that for some subjects, judicial citation means squat compared to the law on the books – statutory law. The U.S. Tax Court is exactly one of these contexts, and we set out to analyze the statutory citation behavior of the U.S. Tax Court from 1990 to 2008. Parsing and analyzing the Tax Court written decisions was probably one of the most interesting and challenging projects I’ve worked on. However, I don’t want to spoil too much. If you’re interested, go grab a copy of the journal from the Virginia Tax Review or download a copy from SSRN. The abstract is below: What can empirical data tell us about the jurisprudence of United States Tax Court? Which sections of the Internal Revenue Code are most frequently cited and has recent tax legislation sparked change in the Tax Court’s decisions? This article presents an analysis of the citation practices of the United States Tax Court between 1990 and 2008. While much of the existing empirical legal research focuses on judicial citation, Tax Court decisions overwhelmingly cite statutes rather than case law. Therefore, analysis of the Tax Court must accommodate the primacy of statutory authority in this domain. After providing a brief overview of relevant prior research, we describe the Tax Court’s written decisions in aggregate along a number of dimensions and examine the Tax Court’s citation of the Internal Revenue Code (I.R.C.) in detail. Though in principle Congress writes the tax code on a blank slate and could design an entirely new legal scheme each year, the Tax Court, as expected, overwhelmingly cites provisions that are likely to remain stable from year to year. We provide evidence of changes resulting from major tax legislation and discuss the implications of these findings for subsequent research. Bommarito, Michael James, Katz, Daniel Martin and Isaacs-See, Jillian, An Empirical Survey of the Population of United States Tax Court Written Decisions (March 20, 2011). Virginia Tax Review, Vol. 30, No. 2, 2011. Available at SSRN: http://ssrn.com/abstract=1441007 ### Featured in Wired: Measuring the Complexity of the Law **Date:** 2014-06-03 | **Category:** Research Sam Arbesman featured our paper on measuring legal complexity of the United States Code on his Wired Science blog, the Social Dimension. Thanks to Sam Arbesman (@arbesman) for featuring Dan and my paper, Measuring the Complexity of the Law: The United States Code, on his excellent Wired Science blog, the Social Dimension. You can read the article here: http://www.wired.com/2014/06/scienceblogs0602law/ and, as a reminder, all of the code and data from the paper is available in this github repository: https://github.com/mjbommar/us-code-complexity Abstract below: Einstein’s razor, a corollary of Ockham’s razor, is often paraphrased as follows: make everything as simple as possible, but not simpler. This rule of thumb describes the challenge that designers of a legal system face — to craft simple laws that produce desired ends, but not to pursue simplicity so far as to undermine those ends. Complexity, simplicity’s inverse, taxes cognition and increases the likelihood of suboptimal decisions. In addition, unnecessary legal complexity can drive a misallocation of human capital toward comprehending and complying with legal rules and away from other productive ends. While many scholars have offered descriptive accounts or theoretical models of legal complexity, empirical research to date has been limited to simple measures of size, such as the number of pages in a bill. No extant research rigorously applies a meaningful model to real data. As a consequence, we have no reliable means to determine whether a new bill, regulation, order, or precedent substantially effects legal complexity. In this paper, we address this need by developing a proposed empirical framework for measuring relative legal complexity. This framework is based on “knowledge acquisition,” an approach at the intersection of psychology and computer science, which can take into account the structure, language, and interdependence of law. We then demonstrate the descriptive value of this framework by applying it to the U.S. Code’s Titles, scoring and ranking them by their relative complexity. Our framework is flexible, intuitive, and transparent, and we offer this approach as a first step in developing a practical methodology for assessing legal complexity. ### Two New Papers on SSRN: Measuring EU Integration Through Sovereign Debt & Exploring Relationships Between Headnotes in the Supreme Court **Date:** 2011-04-18 | **Category:** Research Announcing two working papers released on SSRN: one measuring European integration through sovereign bond yield correlations from 1872 to 2010, and another exploring the network structure of legal concepts in Supreme Court headnotes. What do you do with that unfinished paper? You know, the one that's 50% there but you don't have the time to finish. Or maybe it's the one that's 80% there, but you don't want to deal with the inevitable two years of revise-and-resubmit. This problem gets even harder when you decide to leave academia. I still want to collaborate and carry out public, academic research; however, the amount of time I can dedicate to these projects is now significantly constrained. So what do you do with the projects and papers that didn't make the cut? Today, I decided that I'd rather set some of these papers free on SSRN than let them forever sit on a backup drive. The first paper is an EU integration paper I wrote for a Ph.D. seminar last fall on International Political Economy with Andrew Kerner. It's decent enough, but I don't want to deal with the political economy journal culture. The second is an old SCOTUS project on a network analysis of Lexis headnotes that never got off the ground. Both are interesting, original, and contain enough finished content that I thought they were worth posting. Abstracts and links below: M.J. Bommarito II. The Race to the Bund: Aggregate European and State-Level Integration From the EEC to Present, 2010. Many scholars point to the decreasing yield spread between the benchmark German bund and other European states as a sign of unprecedented European integration since Maastricht. There are two possible issues with this claim. First, without a historical perspective on sovereign bond yields and their comovements, researchers cannot determine how current levels of integration compare to previous periods of European history. To my knowledge, no previous research examines European yields from the formation of the ECSC or EEC to present, let alone pre-war spreads. Second, the first two moments of yield spreads do not necessarily capture meaningful integration. While the spreads are certainly an important measure, we can better assess integration by examining the correlation of these yields. In this article, I address both of these issues as follows. First, I describe an alternative method of measuring integration in the sovereign bond market based on an eigendecomposition of the time-dependent correlation of yields and their first-order differences. This approach better measures integration than standard methods because it takes into account the degree to which a single dimension explains the structure and movement of yields. Second, I construct a dataset of ten-year sovereign bond yields for the EU-15 countries from 1987 to 2010, as well as subsets of these 15 countries from 1958-2010 and from 1872-1913. I then apply both standard yield spread analysis and the method proposed in this article to these three samples. The results of the standard yield spread analysis indicate that Europe has seen previous periods of comparable financial integration. However, based on the alternative method of this article, integration did not occur until the Single European Act of 1987 and the Maastricht Treaty of 1993. A period of stable and high integration is observed following the adoption of the Euro in 1999. Furthermore, both methods show that the recent global real estate and European fiscal crises have significantly degraded financial integration to a level not seen since Maastricht. In summary, this article contributes both a novel and useful methodological approach and a historical basis on which to evaluate modern European integration in the sovereign bond market. M.J. Bommarito II. Exploring Relationships between Legal Concepts in the United States Supreme Court, 2009. We empirically investigate the properties of concepts within the United States Supreme Court corpus using LexisNexis headnotes. By analyzing the relationships between concepts, we uncover structure embedded both within cases and across cases. Furthermore, we estimate the dimensionality of the Supreme Court concept space and thus identify the dominant concepts of the corpus both within a given period and cumulatively. These results provide insight both into the history of these legal concepts and the Court as a whole, proving valuable to a wide range of legal scholars. ### Is There Tax in the Cloud? **Date:** 2014-05-19 | **Category:** Archive A discussion of the U.S. tax implications of cloud computing transactions, highlighting Orly Mazur's SSRN paper on the challenges that IaaS, PaaS, and SaaS create for federal income tax principles. Do you contract for IaaS, PaaS, or SaaS services like Amazon Web Services or SalesForce? Do you provide SaaS services like "cloud-based applications" or "web applications?" What state are you headquartered in? What states are your servers located in? What states are your customers in, even if for just a moment? Depending on your answer to the questions above, you might need to think twice before processing or generating that invoice. The article with abstract below, Taxing the Cloud, by Orly Mazur, is a must-read for your corporate counsel to understand tax in the cloud. Transacting business in the "cloud" has quickly gained popularity worldwide as the new method of providing information technology resources. Instead of purchasing or downloading software, we can now use the Internet to access software and other fundamental computing resources located on remote computer networks operated by third parties. These transactions offer companies lower operating costs, increased scalability and improved reliability, but also give rise to a host of international tax issues. Despite the rapid growth and prevalent use of cloud computing, U.S. taxation of international cloud computing transactions has yet to receive significant scholarly attention. This Article seeks to fill that void by analyzing the U.S. tax implications of operating in the cloud from a doctrinal and policy perspective. Such an analysis shows that the technological advances associated with the cloud put pressure on traditional U.S. federal income tax principles, which creates uncertainty, compliance burdens and liability risks for companies and a potential loss of revenue for the government. Applying the current law to cloud computing transactions also results in tax consequences that run counter to sound tax policy and may result in double taxation or complete non-taxation of cloud income. In light of these problems, federal attention is warranted to clarify how U.S. federal income tax principles apply to businesses operating in the cloud. Thus, this Article proposes that Treasury issue guidance that clearly addresses the U.S. tax implications of international cloud computing services and suggests that, ultimately, the United States must collaborate with other countries to achieve international consensus on these issues. Together these changes will ensure that the United States appropriately taxes the cloud and does so in a manner that minimizes double taxation and promotes efficiency, equity and administrative simplicity. Mazur, Orly, Taxing the Cloud (April 2, 2014). Available at SSRN: http://ssrn.com/abstract=2419644 or http://dx.doi.org/10.2139/ssrn.2419644 ### Law's Future from Finance's Past: Recorded Talk from Reinvent Law Silicon Valley **Date:** 2013-05-18 | **Category:** Archive The recorded video of Michael Bommarito's talk at the Reinvent Law Silicon Valley event, drawing parallels between the evolution of finance and the future trajectory of the legal industry. Back in March, I posted the slides to my talk at the Silicon Valley Reinvent Law event – Law's Future from Finance's Past. Last week, we posted the video online; you can watch below. Michael Bommarito – Law’s Future from Finance’s Past from ReInvent Law Channel on Vimeo: http://vimeo.com/65836938 --- ## Contact Email: inquiry@bommaritollc.com We typically respond within one business day. Initial consultations are complimentary. All inquiries are handled with discretion. We can help with: AI strategy, board advisory, expert testimony, computational modeling, governance & risk, speaking engagements, privacy compliance, and AI auditing. --- ## Privacy This website is a static site that does not use cookies, tracking scripts, or analytics services that collect personal information. Any information provided via email is used solely to respond to inquiries. --- ## Terms Use of this site is governed by terms and disclaimers. Content is provided for educational and informational purposes only and does not constitute legal, financial, or engineering advice.