KonstantinSeurer/mesa

Commit Graph

Author	SHA1	Message	Date
Jason Ekstrand	baf1aa1334	nir: Add foreach_register helper macros	2016-12-29 16:02:44 -08:00
Jason Ekstrand	fb181196de	nir: Rename convert_to_ssa lower_regs_to_ssa This matches the naming of nir_lower_vars_to_ssa, the other to-SSA pass.	2016-12-29 16:02:44 -08:00
Timothy Arceri	194537ebe4	mesa/glsl/i965: remove Driver.NewShader() After removing brw_shader in the previous commit this is no longer needed. V2: remove use in src/compiler/glsl/test_optpass.cpp Reviewed-by: Eric Anholt <eric@anholt.net>	2016-12-30 10:57:17 +11:00
Timothy Arceri	718a0cf49f	i965: move compiled_once flag to brw_program This allows us to delete brw_shader and removes the last use of gl_linked_shader in the codegen paths. Reviewed-by: Eric Anholt <eric@anholt.net>	2016-12-30 10:57:16 +11:00
Timothy Arceri	8417bf528e	mesa/glsl: move BlendSupport bitfield to gl_program This will let us to make _CurrentFragmentProgram a gl_program pointer allowing for simpilifications to be made. We also need to add a field to gl_shader to hold it during parsing. In gl_program we put it inside a union in anticipation of moving more fields here that can be only fs or vertex stage fields. Reviewed-by: Eric Anholt <eric@anholt.net>	2016-12-30 10:57:16 +11:00
Timothy Arceri	3177eef392	mesa: store gl_program in gl_transform_feedback_object rather than gl_shader_program This will allow us to make the CurrentProgram array store gl_program which allows us to do a bunch of simplifications. Reviewed-by: Eric Anholt <eric@anholt.net>	2016-12-30 10:57:16 +11:00
Timothy Arceri	700bc94dce	mesa/glsl: move LinkedTransformFeedback from gl_shader_program to gl_program This will help allow us to store gl_program in the CurrentProgram array rather than gl_shader_program which will allow a bunch of simplifications. Note that we make LinkedTransformFeedback a pointer so we don't waste memory creating a struct for each stage. We also store a pointer to the gl_program that will contain the pointer in gl_shader_program so we can get easy access to the correct stage. Reviewed-by: Eric Anholt <eric@anholt.net>	2016-12-30 10:57:16 +11:00
Timothy Arceri	31c04e4e22	i965: get LinkedTransformFeedback from gl_transform_feedback_object We have already set the gl_shader_program pointer to the correct shader program in _mesa_BeginTransformFeedback() so use it. This is more consistent with how we do it for gen7. Reviewed-by: Eric Anholt <eric@anholt.net>	2016-12-30 10:57:16 +11:00
Timothy Arceri	29d70f5de9	mesa: move _Used to gl_program We no longer need to initialise it because gl_program is never reused. Reviewed-by: Eric Anholt <eric@anholt.net>	2016-12-30 10:57:16 +11:00
Timothy Arceri	8a69ae5345	mesa/compiler: add local_size_variable to shader_info This will be used in api_validate.c in a following patch when we switch to using gl_program pointers for the pipelines CurrentProgram array. Reviewed-by: Eric Anholt <eric@anholt.net>	2016-12-30 10:57:16 +11:00
Timothy Arceri	9ea513e226	mesa: pass gl_program to _mesa_append_uniforms_to_file() This now contains everything we need. Reviewed-by: Eric Anholt <eric@anholt.net>	2016-12-30 10:57:16 +11:00
Timothy Arceri	b51bfbdd85	glsl/mesa: set separate_shader directly in shader_info Reviewed-by: Eric Anholt <eric@anholt.net>	2016-12-30 10:57:16 +11:00
Timothy Arceri	41dd6c3539	mesa/glsl: move subroutine metadata to gl_program This will allow us to store gl_program rather than gl_shader_program as the current program perstage which allows us to simplify code that makes use of the CurrentProgram list. Reviewed-by: Eric Anholt <eric@anholt.net>	2016-12-30 10:57:16 +11:00
Timothy Arceri	0de6f6223a	mesa/compiler: add stage to shader_info This will allow us to simplify the current program logic for SSO. Also since we aim to detach shader_info from nir_shader this will come in handy avoiding passing nir_shader around just to keep track of the stage we are dealing with. V2: set stage for arb asm programs also. Reviewed-by: Eric Anholt <eric@anholt.net>	2016-12-30 10:57:15 +11:00
Eric Anholt	88b41239f9	vc4: Rework scheduling of thread switch to cut one more NOP. Jonas's patch got us most of the benefit of scheduling instructions into the delay slots of thread switch, but if there had been nothing to pair the thrsw with, it would move the thrsw up and leave a NOP where the thrsw was. Instead, don't pair anything with thrsw through the normal scheduling path, and have a separate helper function that inserts the thrsw earlier if possible and inserts any necessary NOPs. total instructions in shared programs: 93027 -> 92643 (-0.41%) instructions in affected programs: 14952 -> 14568 (-2.57%)	2016-12-29 15:22:54 -08:00
Jonas Pfeil	d82dbc4cde	vc4: Fill thread switching delay slots Scan for instructions without a signal set in front of the switching instruction and move the signal up there. shader-db results: total instructions in shared programs: 94494 -> 93027 (-1.55%) instructions in affected programs: 23545 -> 22078 (-6.23%) v2: Fix re-emitting of the instruction in the loop trying to emit NOPs, drop a scheduling change from branch delay slots. (by anholt) Signed-off-by: Jonas Pfeil <pfeiljonas@gmx.de>	2016-12-29 14:41:09 -08:00
Eric Anholt	63e7671c7e	vc4: Enable NIR-based loop unrolling. This successfully unrolls a new shader in GLB2.7, which also gets that shader to successfully compile in multithreaded mode.	2016-12-29 14:41:09 -08:00
Timothy Arceri	5f323198ea	nir: stop gcc warning about uninitialised variables Reviewed-by: Jason Ekstrand <jason@jlekstrand.net>	2016-12-29 13:47:11 +11:00
Dave Airlie	44f833ab18	radv: denote support for extended storage image formats. I'm sure anv has support for these as well, but this is just a first use of the interface to allow different supported spir-v features. Reviewed-by: Bas Nieuwenhuizen <bas@basnieuwenhuizen.nl> Signed-off-by: Dave Airlie <airlied@redhat.com>	2016-12-28 22:44:40 +00:00
Dave Airlie	de7dd4d621	spirv: add interface for drivers to define support extensions. I expect over time the struct contents will change as all drivers support stuff etc, but for now this should be a good starting point. Reviewed-by: Edward O'Callaghan <funfunctor@folklore1984.net> Reviewed-by: Bas Nieuwenhuizen <bas@basnieuwenhuizen.nl> Acked-by: Jason Ekstrand <jason@jlekstrand.net> Signed-off-by: Dave Airlie <airlied@redhat.com>	2016-12-28 22:43:17 +00:00
Chad Versace	464b23b1f2	mesa/shaderobj: Fix races on refcounts Use atomic ops when updating gl_shader::RefCount. Fixes intermittent failures and crashes in 'dEQP-EGL.functional.sharing.gles2.multithread.*'. All tests in that group now pass except 'dEQP-EGL.functional.sharing.gles2.multithread.simple_egl_server_sync.textures.copyteximage2d_texsubimage2d_render'. Tested with: mesa: branch 'master' at `d6545f2` deqp: branch 'nougat-cts-dev' at 4acf725 with additional local fixes DEQP_TARGET: x11_egl hw: Intel Broadwell 0x1616 Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=99085 Reviewed-by: Timothy Arceri <timothy.arceri@collabora.com> Reviewed-by: Tapani Pälli <tapani.palli@intel.com> Cc: mesa-stable@lists.freedesktop.org Cc: Mark Janes <mark.a.janes@intel.com> Cc: Haixia Shi <hshi@chromium.org>	2016-12-28 11:10:43 -08:00
Rob Clark	ec01ef2db1	freedreno/ir3: fix linkage::var size It should actually be 32 for a4xx/a5xx.. we still only advertise 16 but for a5xx the linkage map includes position/psize. Signed-off-by: Rob Clark <robdclark@gmail.com>	2016-12-27 16:54:01 -05:00
Rob Clark	c416ea31cf	freedreno/ir3: treat clipvertex like a normal varying We need this in case it is streamed out. Not sure why we were treating it specially before. Having it as a VS out is harmless if FS doesn't have a matching input. Signed-off-by: Rob Clark <robdclark@gmail.com>	2016-12-27 16:54:01 -05:00
Rob Clark	d10c5a2481	freedreno/a5xx: transform-feedback support We'll need to revisit when adding hw binning pass support, whether we can still do this in main draw step, as we do w/ a3xx/a4xx, or if we needed to move it to the binning stage. Still some failing piglits but most tests pass and the common cases seem to work. Signed-off-by: Rob Clark <robdclark@gmail.com>	2016-12-27 16:54:01 -05:00
Rob Clark	928e9bd602	freedreno: update generated headers Pull in a5xx streamout related regs. Also fixes a couple incorrect register definitions. Signed-off-by: Rob Clark <robdclark@gmail.com>	2016-12-27 16:54:01 -05:00
Rob Clark	6d77ceb701	freedreno/ir3: UBO support for 64b GPUs (a5xx) Update address calculation to support 64b addresses. Signed-off-by: Rob Clark <robdclark@gmail.com>	2016-12-27 16:54:01 -05:00
Rob Clark	fc10dc9fde	freedreno/ir3: rework location of driver constants Rework how we lay out driver constants (driver-params, UBO/TFBO buffer addresses, immediates) for more flexibility. For a5xx+ we need to deal with the fact that gpu ptrs are 64b instead of 32b, which makes the fixed offset scheme not work so well. While we are dealing with that we might also make the layout more dynamic to account for varying # of UBOs, etc. Signed-off-by: Rob Clark <robdclark@gmail.com>	2016-12-27 16:54:01 -05:00
Rob Clark	09202cde7e	freedreno/a5xx: fix emit for bo addresses Reloc for the buffer address is two dwords on 64b devices (a5xx+) Signed-off-by: Rob Clark <robdclark@gmail.com>	2016-12-27 16:54:01 -05:00
Rob Clark	f043904080	freedreno/a5xx: texture layout Seems to be imilar to a4xx, and sampler state "array-pitch" needs to be aligned to page size. Signed-off-by: Rob Clark <robdclark@gmail.com>	2016-12-27 16:54:01 -05:00
Rob Clark	859cb24d94	ttn: set ->info->num_ubos For dealing w/ 32b vs 64b gpu addresses, I need to rework how we pass UBO buffer addresses to shader, and knowing up front the # of UBOs is useful. But I noticed ttn wasn't setting this. Signed-off-by: Rob Clark <robdclark@gmail.com> Reviewed-by: Eric Anholt <eric@anholt.net>	2016-12-27 16:54:01 -05:00
Chad Versace	d6545f2345	anv: Handle vkGetPhysicalDeviceQueueFamilyProperties with count == 0 The spec implicitly allows the incoming count to be 0. From the Vulkan 1.0.38 spec, Section 4.1 Physical Devices: If the value referenced by pQueueFamilyPropertyCount is not 0 [then do stuff]. Cc: mesa-stable@lists.freedesktop.org Reviewed-by: Anuj Phogat <anuj.phogat@gmail.com> Reviewed-by: Jason Ekstrand <jason@jlekstrand.net>	2016-12-27 12:31:34 -08:00
Chad Versace	b85c0b569f	egl: Emit correct error when robust context creation fails Fixes dEQP-EGL.functional.create_context_ext.robust_* on Intel with GBM. If the user sets the EGL_CONTEXT_OPENGL_ROBUST_ACCESS_BIT_KHR in EGL_CONTEXT_FLAGS_KHR when creating an OpenGL ES context, then EGL_KHR_create_context spec requires that we unconditionally emit EGL_BAD_ATTRIBUTE because that flag does not exist for OpenGL ES. When creating an OpenGL context, the spec requires that we emit EGL_BAD_MATCH if we can't support the request; that error is generated in the egl_dri2 layer where the driver capability is actually checked. Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=99188 Cc: mesa-stable@lists.freedesktop.org Reviewed-by: Tapani Pälli <tapani.palli@intel.com>	2016-12-27 10:21:29 -08:00
Damien Grassart	75252826e8	anv: return count of queue families written The Vulkan spec indicates that vkGetPhysicalDeviceQueueFamilyProperties() should overwrite pQueueFamilyPropertyCount with the number of structures actually written to pQueueFamilyProperties. Signed-off-by: Damien Grassart <damien@grassart.com> Reviewed-by: Chad Versace <chadversary@chromium.org> Cc: mesa-stable@lists.freedesktop.org	2016-12-27 10:15:47 -08:00
Chad Versace	e2d69d5e2d	i965: Allow import/export of ARGB1555 images To my knowledge, this fixes no tests. I simply wrote the patch for completeness as a follow-up to the previous two patches. Reviewed-by: Tapani Pälli <tapani.palli@intel.com>	2016-12-27 09:14:04 -08:00
Chad Versace	f3739810e3	mesa/texformat: Handle GL_RGBA + GL_UNSIGNED_SHORT_5_5_5_1 _mesa_choose_tex_format() already handles GL_RGBA + GL_UNSIGNED_SHORT_1_5_5_5_REV by converting it to MESA_FORMAT_B5G5R5A1_UNORM. Teach it do the same for the non-reversed type. Otherwise, the switch's fallthrough converts it to an 8888 format, which has incompatible precision in the alpha channel. Patch 2/2 to fix dEQP-EGL.functional.image.modify.tex_rgb5_a1_tex_subimage_rgba8 on Intel. Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=99185 Cc: Haixia Shi <hshi@chromium.org> Reviewed-by: Tapani Pälli <tapani.palli@intel.com> Cc: "13.0" <mesa-stable@lists.freedesktop.org>	2016-12-27 09:14:00 -08:00
Chad Versace	9aa6ab0748	dri: Add __DRI_IMAGE_FORMAT_ARGB1555 This allows eglCreateImage() to accept textures of said format. Patch 1/2 to fix dEQP-EGL.functional.image.modify.tex_rgb5_a1_tex_subimage_rgba8 on Intel. Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=99185 Cc: Haixia Shi <hshi@chromium.org> Reviewed-by: Tapani Pälli <tapani.palli@intel.com> Cc: "13.0" <mesa-stable@lists.freedesktop.org>	2016-12-27 09:13:43 -08:00
Tapani Pälli	4d6d4f939e	egl/dri2: implement query surface hook This makes better guarantee that the values we return are in sync what the underlying drawable currently has. Together with dEQP change in bug #98327 this fixes following test: dEQP-EGL.functional.resize.surface_size.grow v2: avoid unnecessary x11 roundtrips (Chad Versace) Signed-off-by: Tapani Pälli <tapani.palli@intel.com> Tested-by: Mark Janes <mark.a.janes@intel.com> Reviewed-by: Chad Versace <chadversary@chromium.org> Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=98327	2016-12-27 08:01:08 +02:00
Dave Airlie	d8423772ca	radv: add some asserts for operations on general queue These might be useful in the future, or not. Signed-off-by: Dave Airlie <airlied@redhat.com>	2016-12-27 03:27:14 +00:00
Bas Nieuwenhuizen	059af2515a	radv: Also skip DCC clear flushes for compute. (airlied: fixes DOOM hang with compute queue enabled) Reviewed-by: Dave Airlie <airlied@redhat.com> Signed-off-by: Bas Nieuwenhuizen <basni@google.com>	2016-12-27 03:27:13 +00:00
Dave Airlie	3fd306b423	radv: handle queue present directly to winsys Don't call the QueueSubmit interface, just call direct to the winsys, so we can pass the wait semaphores. Noticed while debugging doom, doesn't fix anything. Reviewed-by: Bas Nieuwenhuizen <bas@basnieuwenhuizen.nl> Signed-off-by: Dave Airlie <airlied@redhat.com>	2016-12-26 22:20:35 +00:00
Jordan Justen	097c9dc2d4	intel/blorp_blit: Fix max blit size for gen6 Fixes ES3-CTS.gtf.GL3Tests.framebuffer_blit.framebuffer_blit_functionality_stencil_blit Signed-off-by: Jordan Justen <jordan.l.justen@intel.com> Reviewed-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com> Reviewed-by: Kenneth Graunke <kenneth@whitecape.org>	2016-12-26 08:50:21 -08:00
Dave Airlie	b5bb8b54cf	radv: fix rendering to b10g11r11_ufloat_pack32 doom was causing a printf about an illegal color, it was due the non-void returning -1, and the other function checking for 4, align these. Reviewed-by: Edward O'Callaghan <funfunctor@folklore1984.net> Reviewed-by: Bas Nieuwenhuizen <bas@basnieuwenhuizen.nl> Signed-off-by: Dave Airlie <airlied@redhat.com>	2016-12-26 10:31:20 +10:00
Dave Airlie	4813c9ade7	radv: handle multi-component shared load/stores. This was seen in doom shaders, so handle it properly. Reviewed-by: Bas Nieuwenhuizen <bas@basnieuwenhuizen.nl> Signed-off-by: Dave AIrlie <airlied@redhat.com>	2016-12-26 10:31:20 +10:00
Vedran Miletić	d9fef848a6	clover: Use Clang's diagnostics Presently errors from frontend are handled only if they occur in clang::CompilerInvocation::CreateFromArgs(). This patch uses clang::DiagnosticsEngine to detect errors such as invalid values for Clang frontend arguments. Fixes Piglit's cl/program/build/fail/invalid-version-declaration.cl test. v2: fix inconsistent code formatting Signed-off-by: Vedran Miletić <vedran@miletic.net> Reviewed-by: Francisco Jerez <currojerez@riseup.net> Tested-by: Aaron Watry <awatry@gmail.com>	2016-12-24 18:35:09 -08:00
Damien Grassart	3a30b1a556	radv: return count of queue families written The Vulkan spec indicates that vkGetPhysicalDeviceQueueFamilyProperties() should overwrite pQueueFamilyPropertyCount with the number of structures actually written to pQueueFamilyProperties. Signed-off-by: Damien Grassart <damien@grassart.com> Reviewed-by: Bas Nieuwenhuizen <bas@basnieuwenhuizen.nl>	2016-12-25 02:25:02 +01:00
Jason Ekstrand	88b5acfa09	i965/generator/tex: Handle an immediate sampler with an indirect texture In this case we were dying when we tried to do SHL addr sampler imm(8) because that puts an immediate in src0 of a two source instruction. This fixes 2704 of the new separate sampler Vulkan CTS tests on Sky Lake. Reviewed-by: Eduardo Lima Mitev <elima@igalia.com> Cc: "13.0" <mesa-stable@lists.freedesktop.org>	2016-12-23 07:27:13 -08:00
Bruce Cherniak	9e35426731	swr: fix icc compile error ICC doesn't like the use of nullptr (std::nullptr_t) argument in p_atomic_set. GCC and clang don't complain. Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=99119 Reviewed-by: Tim Rowley <timothy.o.rowley@intel.com>	2016-12-23 08:36:21 -06:00
Dave Airlie	e7279f16a0	radv: set some proper values for interp offset limits. These are taken from the amdgpu-pro driver, and cause no CTS change. Reviewed-by: Bas Nieuwenhuizen <bas@basnieuwenhuizen.nl> Signed-off-by: Dave Airlie <airlied@redhat.com>	2016-12-23 14:36:54 +10:00
Dave Airlie	14737bcdd5	radv: bump texel offsets to align with radeonsi it appears from the amdgpu-pro results the hw can do more, but let's just align with radeonsi for now. No CTS regressions. Reviewed-by: Bas Nieuwenhuizen <bas@basnieuwenhuizen.nl> Signed-off-by: Dave Airlie <airlied@redhat.com>	2016-12-23 14:36:50 +10:00
Jason Ekstrand	d55835b8bd	nir/algebraic: Add optimizations for "a == a && a CMP b" This sequence shows up The Talos Principal, at least under Vulkan, and prevents loop analysis from properly computing trip counts in a few loops. Reviewed-by: Ian Romanick <ian.d.romanick@intel.com>	2016-12-22 16:27:19 -08:00
Jason Ekstrand	8962cc96ec	i965: Use nir_opt_trivial_continues and nir_opt_if Reviewed-by: Timothy Arceri <timothy.arceri@collabora.com>	2016-12-22 16:27:19 -08:00
Jason Ekstrand	6d9f576b56	nir: Add a pass for moving SPIR-V continue blocks to the ends of loops When shaders come in from SPIR-V, we handle continue blocks by placing the contents of the continue inside of a "if (!first_iteration)". We do this so that we can properly handle the fact that continues in SPIR-V jump to the continue block at the end of the loop rather than jumping directly to the top of the loop like they do in NIR. In particular, the increment step of a simple for loop ends up in the continue block. This pass looks for this case in loops that don't actually have any continues and moves the continue contents to the end of the loop instead. We need this because loop unrolling doesn't work if the increment is inside of a condition. Reviewed-by: Timothy Arceri <timothy.arceri@collabora.com>	2016-12-22 16:27:19 -08:00
Jason Ekstrand	1111a05f90	nir: Add an optimization pass to remove trivial continues Reviewed-by: Timothy Arceri <timothy.arceri@collabora.com>	2016-12-22 16:27:19 -08:00
Jason Ekstrand	993e9195d4	nir: Correctly handle blocks in cf_node_cf_tree_next Reviewed-by: Timothy Arceri <timothy.arceri@collabora.com>	2016-12-22 16:27:19 -08:00
Timothy Arceri	3321eb4c36	i965: make use of nir_lower_returns() for GL Fixes two new piglit tests: spec/glsl-1.10/execution/vs-nested-return-sibling-loop.shader_test spec/glsl-1.10/execution/vs-nested-return-sibling-loop2.shader_test shader-db results for BDW: total instructions in shared programs: 12903158 -> 12903134 (-0.00%) instructions in affected programs: 27100 -> 27076 (-0.09%) helped: 32 HURT: 6 total cycles in shared programs: 294922518 -> 294922804 (0.00%) cycles in affected programs: 4372828 -> 4373114 (0.01%) helped: 31 HURT: 8 Reviewed-by: Kenneth Graunke <kenneth@whitecape.org>	2016-12-23 10:59:32 +11:00
Timothy Arceri	f20ba7ad44	nir: update nir_lower_returns to only predicate instructions when needed Unless an if statement contains nested returns we can simply add any following instructions to the branch without the return. V2: fix handling if_nested_return value when there is a sibling if/loop that doesn't contain a return. (Spotted by Ken) V3: - add a better comment to the new variable - remove instructions after if when both branches return Reviewed-by: Jason Ekstrand <jason@jlekstrand.net>	2016-12-23 10:59:32 +11:00
Timothy Arceri	40e9f2f138	i965: disable loop unrolling in GLSL IR There is a single regression in loop unrolling which is: loops HURT: shaders/orbital_explorer.shader_test GS SIMD8: 0 -> 1 However the loop is huge so it seems reasonable not to unroll it. It's surprising that GLSL IR does unroll it. shader-db results BDW: total instructions in shared programs: 13037455 -> 13036947 (-0.00%) instructions in affected programs: 17982 -> 17474 (-2.83%) helped: 63 HURT: 25 total cycles in shared programs: 262217870 -> 262227990 (0.00%) cycles in affected programs: 2287046 -> 2297166 (0.44%) helped: 969 HURT: 844 total loops in shared programs: 2951 -> 2952 (0.03%) loops in affected programs: 0 -> 1 helped: 0 HURT: 1 LOST: 0 GAINED: 1 Reviewed-by: Jason Ekstrand <jason@jlekstrand.net>	2016-12-23 10:15:36 +11:00
Timothy Arceri	715f0d06d1	i965: use nir loop unrolling pass shader-db results for BDW: total instructions in shared programs: 12589614 -> 12590119 (0.00%) instructions in affected programs: 50525 -> 51030 (1.00%) helped: 7 HURT: 145 total cycles in shared programs: 241524604 -> 241490502 (-0.01%) cycles in affected programs: 1941404 -> 1907302 (-1.76%) helped: 302 HURT: 449 total loops in shared programs: 4245 -> 2947 (-30.58%) loops in affected programs: 1535 -> 237 (-84.56%) helped: 1142 HURT: 0 total spills in shared programs: 14453 -> 14453 (0.00%) spills in affected programs: 0 -> 0 helped: 0 HURT: 0 total fills in shared programs: 18984 -> 18984 (0.00%) fills in affected programs: 0 -> 0 helped: 0 HURT: 0 LOST: 26 GAINED: 15 Reviewed-by: Jason Ekstrand <jason@jlekstrand.net>	2016-12-23 10:15:36 +11:00
Timothy Arceri	e729504fb1	nir: pass compiler rather than devinfo to functions that call nir_optimize Later we will pass compiler to nir_optimise to be used by the loop unroll pass. Reviewed-by: Jason Ekstrand <jason@jlekstrand.net>	2016-12-23 10:15:36 +11:00
Timothy Arceri	51daccb289	nir: add a loop unrolling pass V2: - tidy ups suggested by Connor. - tidy up cloning logic and handle copy propagation based of suggestion by Connor. - use nir_ssa_def_rewrite_uses to fix up lcssa phis suggested by Connor. - add support for complex loop unrolling (two terminators) - handle case were the ssa defs use outside the loop is already a phi - support unrolling loops with multiple terminators when trip count is know for each terminator V3: - set correct num_components when creating phi in complex unroll - rewrite update remap table based on Jasons suggestions. - remove unrequired extract_loop_body() helper as suggested by Jason. - simplify the lcssa phi fix up code for simple loops as per Jasons suggestions. - use mem context to keep track of hash table memory as suggested by Jason. - move is_{complex,simple}_loop helpers to the unroll code - require nir_metadata_block_index - partially rewrote complex unroll to be simpler and easier to follow. V4: - use rzalloc() when creating nir_phi_src but not setting pred right away fixes regression cause by ralloc() no longer zeroing memory. V5: - simplify calling of complex_unroll() - use new loop terminator fields to get the break/continue from blocks and simplify loop unrolling code - handle slightly less trivial loop terminators. if branches can now have instructions but can only contain a single block. - use nir print type IR snippets in unroll function descriptions - add better explanation and variable for why we need to clone additional times when the second terminator it the limiting terminator. - partially convert out of ssa before unrolling loops (suggested by Jason) v6: - remove unused nir_builder - use Jasons new from ssa helper - tidy/fixup cursor use - unroll terminators that contain control flow correctly - unroll complex loops with control flow before the terminators correctly Reviewed-by: Jason Ekstrand <jason@jlekstrand.net>	2016-12-23 10:15:36 +11:00
Timothy Arceri	f8407a5398	nir: add helper for cloning nir_cf_list V2: - updated to create a generic list clone helper nir_cf_list_clone() - continue to assert on clone when fallback flag not set as suggested by Jason. Reviewed-by: Jason Ekstrand <jason@jlekstrand.net>	2016-12-23 10:15:36 +11:00
Timothy Arceri	b84dfa0f62	nir: update fixup_phi_srcs() to handle registers We need to do this because we partially get out of SSA when unrolling and cloning loops. Reviewed-by: Jason Ekstrand <jason@jlekstrand.net>	2016-12-23 10:15:36 +11:00
Timothy Arceri	d781320974	nir: create helper for fixing phi srcs when cloning This will be useful for fixing phi srcs when cloning a loop body during loop unrolling. Reviewed-by: Jason Ekstrand <jason@jlekstrand.net>	2016-12-23 10:15:36 +11:00
Thomas Helland	ec8423a4b1	nir: Add a LCSAA-pass V2: Do a "depth first search" to convert to LCSSA V3: Small comment fixup V4: Rebase, adapt to removal of function overloads V5: Rebase, adapt to relocation of nir to compiler/nir Still need to adapt to potential if-uses Work around nir_validate issue V6 (Timothy): - tidy lcssa and stop leaking memory - dont rewrite the src for the lcssa phi node - validate lcssa phi srcs to avoid postvalidate assert - don't add new phi if one already exists - more lcssa phi validation fixes - Rather than marking ssa defs inside a loop just mark blocks inside a loop. This is simpler and fixes lcssa for intrinsics which do not have a destination. - don't create LCSSA phis for loops we won't unroll - require loop metadata for lcssa pass - handle case were the ssa defs use outside the loop is already a phi V7: (Timothy) - pass indirect mask to metadata call v8: (Timothy) - make convert to lcssa a helper function rather than a nir pass - replace inside loop bitset with on the fly block index logic. - remove lcssa phi validation special cases - inline code from useless helpers, suggested by Jason. - always do lcssa on loops, suggested by Jason. - stop making lcssa phis special. Add as many source as the block has predecessors, suggested by Jason. V9: (Timothy) - fix regression with the is_lcssa_phi field not being initialised to false now that ralloc() doesn't zero out memory. V10: (Timothy) - remove extra braces in SSA example, pointed out by Topi V11: (Timothy) - add missing support for LCSSA phis in if conditions. V12: (Timothy) - small tidy up suggested by Jason. - always create lcssa phi even if it just points to an lcssa phi from an inner loop Reviewed-by: Jason Ekstrand <jason@jlekstrand.net>	2016-12-23 10:15:36 +11:00
Thomas Helland	6772a17acc	nir: Add a loop analysis pass This pass detects induction variables and calculates the trip count of loops to be used for loop unrolling. V2: Rebase, adapt to removal of function overloads V3: (Timothy Arceri) - don't try to find trip count if loop terminator conditional is a phi - fix trip count for do-while loops - replace conditional type != alu assert with return - disable unrolling of loops with continues - multiple fixes to memory allocation, stop leaking and don't destroy structs we want to use for unrolling. - fix iteration count bugs when induction var not on RHS of condition - add FIXME for && conditions - calculate trip count for unsigned induction/limit vars V4: (Timothy Arceri) - count instructions in a loop - set the limiting_terminator even if we can't find the trip count for all terminators. This is needed for complex unrolling where we handle 2 terminators and the trip count is unknown for one of them. - restruct structs so we don't keep information not required after analysis and remove dead fields. - force unrolling in some cases as per the rules in the GLSL IR pass V5: (Timothy Arceri) - fix metadata mask value 0x10 vs 0x16 V6: (Timothy Arceri) - merge loop_variable and nir_loop_variable structs and lists suggested by Jason - remove induction var hash table and store pointer to induction information in the loop_variable suggested by Jason. - use lowercase list_addtail() suggested by Jason. - tidy up init_loop_block() as per Jasons suggestions. - replace switch with nir_op_infos[alu->op].num_inputs == 2 in is_var_basic_induction_var() as suggested by Jason. - use nir_block_last_instr() in and rename foreach_cf_node_ex_loop() as suggested by Jason. - fix else check for is_trivial_loop_terminator() as per Connors suggetions. - simplify offset for induction valiables incremented before the exit conditions is checked. - replace nir_op_isub check with assert() as it should have been lowered away. V7: (Timothy Arceri) - use rzalloc() on nir_loop struct creation. Worked previously because ralloc() was broken and always zeroed the struct. - fix cf_node_find_loop_jumps() to find jumps when loops contain nested if statements. Code is tidier as a result. V8: (Timothy Arceri) - move is_trivial_loop_terminator() to nir.h so we can use it to assert is the loop unroll pass - fix analysis to not bail when looking for terminator when the break is in the else rather then the if - added new loop terminator fields: break_block, continue_from_block and continue_from_then so we don't have to gather these when doing unrolling. - get correct array length when forcing unrolling of variables indexed arrays that are the same size as the iteration count - add support for induction variables of type float - update trival loop terminator check to allow an if containing instructions as long as both branches contain only a single block. V9: (Timothy) - bunch of tidy ups and simplifications suggested by Jason. - rewrote trivial terminator detection, now the only restriction is there must be no nested jumps, anything else goes. - rewrote the iteration test to use nir_eval_const_opcode(). - count instruction properly even when forcing an unroll. - bunch of other tidy ups and simplifications. V10: (Timothy) - some trivial tidy ups suggested by Jason. - conditional fix for break inside continue branch by Jason. Reviewed-by: Jason Ekstrand <jason@jlekstrand.net>	2016-12-23 10:15:36 +11:00
Timothy Arceri	eda3ec7957	i965: use nir_lower_indirect_derefs() for GLSL This moves the nir_lower_indirect_derefs() call into brw_preprocess_nir() so thats is called by both OpenGL and Vulkan and removes that call to the old GLSL IR pass lower_variable_index_to_cond_assign() We want to do this pass in nir to be able to move loop unrolling to nir. There is a increase of 1-3 instructions in a small number of shaders, and 2 Kerbal Space program shaders that increase by 32 instructions. The changes seem to be caused be the difference in the GLSL IR vs NIR variable index lowering passes. The GLSL IR pass creates a simple if ladder for arrays of size 4 or less, while the NIR pass implements a binary search for all arrays regardless of size. Shader-db results BDW: total instructions in shared programs: 13021176 -> 13021819 (0.00%) instructions in affected programs: 57693 -> 58336 (1.11%) helped: 20 HURT: 190 total cycles in shared programs: 299805580 -> 299750826 (-0.02%) cycles in affected programs: 2290024 -> 2235270 (-2.39%) helped: 337 HURT: 442 total fills in shared programs: 19984 -> 19984 (0.00%) fills in affected programs: 0 -> 0 helped: 0 HURT: 0 LOST: 4 GAINED: 0 V2: remove the do_copy_propagation() call from the i965 GLSL IR linking code. This call was added in `f7741c5211` but since we are moving the variable index lowering to NIR we no longer need it and can just rely on the nir copy propagation pass. Reviewed-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Jason Ekstrand <jason@jlekstrand.net>	2016-12-23 10:15:36 +11:00
Timothy Arceri	976859ce57	i965: allow sampler indirects on all gens Without this we will regress the max-samplers piglit test on Gen6 and lower when loop unrolling is done in NIR. There is a check in the GLSL IR linker that errors when it finds indirects and EmitNoIndirectSampler is set. As far as I can tell there is no reason for not enabling this for all gens regardless of whether they fully support ARB_gpu_shader5 or not. Reviewed-by: Jason Ekstrand <jason@jlekstrand.net>	2016-12-23 10:15:35 +11:00
Jason Ekstrand	a620f66872	nir: Add a couple quick-and-dirty out-of-SSA helpers These are designed for use within an optimization pass when SSA becomes more pain than it's worth. They're very naive and don't generate anything close to optimal register-based NIR. Also, they may result in shaders which do not validate because of, for instance, registers in phi sources. However, the register-based into-SSA pass should be pretty efficient at cleaning up the mess. Reviewed-by: Timothy Arceri <timothy.arceri@collabora.com>	2016-12-23 10:15:35 +11:00
Arda Coskunses	99de7b7525	vulkan/wsi/x11: don't crash on null wsi x11 connection Without this check driver crash when application window closed unexpectedly. Acked-by: Edward O'Callaghan <funfunctor@folklore194.net> Reviewed-by: Jason Ekstrand <jason@jlekstrand.net> Cc: "13.0" <mesa-stable@lists.freedesktop.org>	2016-12-22 14:09:46 -08:00
Arda Coskunses	01dd363e67	vulkan/wsi/x11: don't crash on null visual When application window closed unexpectedly due to lost window visualtypes getting invlaid parameters which is causing a crash. Necessary check is added to prevent the crash. Reviewed-by: Jason Ekstrand <jason@jlekstrand.net> Cc: "13.0" <mesa-stable@lists.freedesktop.org>	2016-12-22 14:09:34 -08:00
Christian Inci	7a4ea95f1c	radeonsi: Bugfix needed for hashcat Hashcat needs MAX_GLOBAL_BUFFERS to be 21 or even 22 for some modes. It'll crash otherwise. I'm adding an assert to see if programs need it to be even higher. Signed-off-by: Christian Inci <chris.bugsfd@broke-the-inter.net> [Handle first properly; should be NFC, since clover always uses first == 0.] Signed-off-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2016-12-22 17:11:43 +01:00
Nicolai Hähnle	eca57f85ee	radeonsi: fix gl_ClipDistance and gl_ClipVertex for points The clipper hardware doesn't consider points as primitives that can be clipped. Simply setting the corresponding cull bits works, and should not have an adverse effect on other primitive types according to the hardware team. Reviewed-by: Marek Olšák <marek.olsak@amd.com> Reviewed-by: Edward O'Callaghan <funfunctor@folklore1984.net>	2016-12-22 16:59:58 +01:00
Nicolai Hähnle	3778a10d37	radeonsi: only set VS_OUT_MISC_SIDE_BUS_ENA when the misc vector is used Should have no effect (other than perhaps on power consumption), but Vulkan does this. Reviewed-by: Marek Olšák <marek.olsak@amd.com> Reviewed-by: Edward O'Callaghan <funfunctor@folklore1984.net>	2016-12-22 16:58:53 +01:00
Vinson Lee	ede8c02ab0	llvmpipe: Link tests with CLOCK_LIB. Fix linking error with 'make check'. CXXLD lp_test_format ../../../../src/gallium/auxiliary/.libs/libgallium.a(os_time.o): In function `os_time_get_nano': src/gallium/auxiliary/os/os_time.c:59: undefined reference to `clock_gettime' Signed-off-by: Vinson Lee <vlee@freedesktop.org>	2016-12-21 17:23:05 -08:00
Fredrik Höglund	27a8aab882	radv: fix dual source blending Add the index to the location when assigning driver locations for output variables. Otherwise two fragment shader outputs declared as: layout (location = 0, index = 0) out vec4 output1; layout (location = 0, index = 1) out vec4 output2; will end up aliasing one another. Note that this patch will make the second output variable in the above example alias a possible third output variable with location = 1 and index = 0. But this shouldn't be a problem in practice since only one color attachment is supported when dual-source blending is used. Cc: "13.0" <mesa-stable@lists.freedesktop.org> Reviewed-by: Bas Nieuwenhuizen <bas@basnieuwenhuizen.nl>	2016-12-22 02:07:17 +01:00
Dave Airlie	877202b6dc	radv: enable shaderStorageImageExtendedFormats This passes all the CTS tests that get enabled for this. Reviewed-by: Bas Nieuwenhuizen <bas@basnieuwenhuizen.nl> Signed-off-by: Dave Airlie <airlied@redhat.com>	2016-12-22 10:29:15 +10:00
Dave Airlie	a3ca2a9b7b	radv: enable shaderGatherImageExtended Thanks to Ilia's patch this works fine on radv. No regressions in CTS, all enabled tests pass. Reviewed-by: Bas Nieuwenhuizen <bas@basnieuwenhuizen.nl> Signed-off-by: Dave Airlie <airlied@redhat.com>	2016-12-22 09:48:18 +10:00
Dave Airlie	56020c7a7c	radv/image: only touch queue family info for concurrent images. The spec says to ignore these fields for exclusive images. Fixes crashes in: dEQP-VK.clipping.* Reviewed-by: Bas Nieuwenhuizen <bas@basnieuwenhuizen.nl> Signed-off-by: Dave Airlie <airlied@redhat.com>	2016-12-21 23:33:04 +00:00
Dave Airlie	9d23b8a18e	radv: flush smem for uniform buffer bit. (cc'ing stable as I'd like to backport the ubo speedup as well) Reviewed-by: Bas Nieuwenhuizen <bas@basnieuwenhuizen.nl> Cc: "13.0" <mesa-stable@lists.freedesktop.org> Signed-off-by: Dave Airlie <airlied@redhat.com>	2016-12-21 22:31:14 +00:00
Junwei Zhang	018ead4266	radeonsi: add Polaris12 support (v3) v2: use gfxip names for llvm 4.0+ v3: use tonga for llvm <= 3.8, drop gfxip name, we can just change that we change the other asics. Reviewed-by: Marek Olšák <marek.olsak@amd.com> Signed-off-by: Junwei Zhang <Jerry.Zhang@amd.com> Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com> Acked-by: Christian König <christian.koenig@amd.com>	2016-12-21 15:10:03 -05:00
Ian Romanick	15c8f322ca	glsl: Eliminate the open-coded version of process_block_array_leaf Signed-off-by: Ian Romanick <ian.d.romanick@intel.com> Reviewed-by: Timothy Arceri <timothy.arceri@collabora.com>	2016-12-21 10:24:45 -08:00
Juan A. Suarez Romero	415f5f09e3	ttn: handle GLSL_SAMPLER_DIM_SUBPASS_MS case Fixes a warning. Reviewed-by: Samuel Iglesias Gonsálvez <siglesias@igalia.com>	2016-12-21 12:44:25 +01:00
Juan A. Suarez Romero	c32a9ec5f5	i965: allow unsourced enabled VAO The GL 4.5 spec says: "If any enabled array’s buffer binding is zero when DrawArrays or one of the other drawing commands defined in section 10.4 is called, the result is undefined." This commits avoids crashing the code, which is not a very good "undefined result". This fixes spec/!opengl 3.1/vao-broken-attrib piglit test.	2016-12-21 12:37:22 +01:00
Edward O'Callaghan	8801734da7	svga: Fix a strict-aliasing violation in shader dumper As per the C spec, it is illegal to alias pointers to different types. This results in undefined behaviour after optimization passes, resulting in very subtle bugs that happen only on a full moon.. Use a memcpy() as a well defined coercion between the isomorphic bit-field interpretations of memory. V.2: Use C99 compat STATIC_ASSERT() over C11 static_assert(). Signed-off-by: Edward O'Callaghan <funfunctor@folklore1984.net> Reviewed-by: Charmaine Lee <charmainel@vmware.com>	2016-12-21 15:00:21 +11:00
Roland Scheidegger	e827d91756	draw: use SoA fetch, not AoS one Now that there's some SoA fetch which never falls back, we should always get results which are better or at least not worse (something like rgba32f will stay the same). For cases which get way better, think something like R16_UNORM with 8-wide vectors: this was 8 sign-extend fetches, 8 cvt, 8 muls, followed by a couple of shuffles to stitch things together (if it is smart enough, 6 unpacks) and then a (8-wide) transpose (not sure if llvm could even optimize the shuffles + transpose, since the 16bit values were actually sign-extended to 128bit before being cast to a float vec, so that would be another 8 unpacks). Now that is just 8 fetches (directly inserted into vector, albeit there's one 128bit insert needed), 1 cvt, 1 mul. v2: ditch the old AoS code instead of just disabling it. Reviewed-by: Jose Fonseca <jfonseca@vmware.com>	2016-12-21 04:48:24 +01:00
Roland Scheidegger	cb81460dcc	gallivm: generalize the compressed format soa fetch a bit This can now handle rgtc (unorm) too - this path no longer handles plain formats, but that's unnecessary they now all have their proper SoA unpack (this will still be dog-slow though due to the actual fetch being per-pixel util fallbacks). Reviewed-by: Jose Fonseca <jfonseca@vmware.com>	2016-12-21 04:48:24 +01:00
Roland Scheidegger	3c98e3cd63	gallivm: provide soa fetch path handling formats with more than 32bit This previously always fell back to AoS conversion. Even for 4-float formats (which is the optimal case by far for that fallback case) this was suboptimal, since it meant the conversion couldn't be done with 256bit vectors. While this may still only be partly possible for some formats, (unless there's AVX2 support) at least the transpose can be done with half the unpacks (and before using the transpose for AoS fallbacks, it was worse still). With less than 4 channels, things got way worse with the AoS fallback quickly even with 128bit vectors. The strategy is pretty much the same as the existing one for formats which fit into 32 bits, except there's now multiple vectors to be fetched (2 or 4 to be exact), which need to be shuffled first (if it's 4 vectors, this amounts to a transpose, for 2 it's a bit different), then the unpack is done the same (with the exception that the shift of the channels is now modulo 32, and we need to select the right vector). In fact the most complex part about it is to get the shuffles right for separating into lo/hi parts for AVX/AVX2... This also makes use of the new ability of gather to use provided type information, which we abuse to outsmart llvm so we get decent shuffles, and to fetch 3x32bit vectors without having to ZExt the scalar. And just because we can, we handle double formats too, albeit they are a bit different (draw sometimes needs to handle that). v2: fix typo float/int bug (generating inefficient code). Reviewed-by: Jose Fonseca <jfonseca@vmware.com>	2016-12-21 04:48:24 +01:00
Roland Scheidegger	8bd67a35c5	gallivm: optimize gather a bit, by using supplied destination type By using a dst_type in the the gather interface, gather has some more knowledge about how values should be fetched. E.g. if this is a 3x32bit fetch and dst_type is 4x32bit vector gather will no longer do a ZExt with a 96bit scalar value to 128bit, but just fetch the 96bit as 3x32bit vector (this is still going to be 2 loads of course, but the loads can be done directly to simd vector that way). Also, we can now do some try to use the right int/float type. This should make no difference really since there's typically no domain transition penalties for such simd loads, however it actually makes a difference since llvm will use different shuffle lowering afterwards so the caller can use this to trick llvm into using sane shuffle afterwards (and yes llvm is really stupid there - nothing against using the shuffle instruction from the correct domain, but not at the cost of doing 3 times more shuffles, the case which actually matters is refusal to use shufps for integer values). Also do some attempt to avoid things which look great on paper but llvm doesn't really handle (e.g. fetching 3-element 8 bit and 16 bit vectors which is simply disastrous - I suspect type legalizer is to blame trying to extend these vectors to 128bit types somehow, so fetching these with scalars like before which is suboptimal due to the ZExt). Remove the ability for truncation (no point, this is gather, not conversion) as it is complex enough already. While here also implement not just the float, but also the 64bit avx2 gathers (disabled though since based on the theoretical numbers the benefit just isn't there at all until Skylake at least). Reviewed-by: Jose Fonseca <jfonseca@vmware.com>	2016-12-21 04:48:24 +01:00
Roland Scheidegger	5b950319ce	gallivm: optimize SoA AoS fallback fetch path a little We should do transpose, not extract/insert, at least with "sufficient" amount of channels (for 4 channels, extract/insert shuffles generated otherwise look truly terrifying). Albeit we shouldn't fallback to that so often in any case. v2: ditch the extract/insert path, not worth keeping (we're going to avoid hitting the fallback that often with future patches). Reviewed-by: Jose Fonseca <jfonseca@vmware.com>	2016-12-21 04:48:24 +01:00
Roland Scheidegger	d7d23aee4b	gallivm: (trivial) handle non-aligned fetch for lp_build_fetch_rgba_soa soa fetch so far always assumed that data was aligned. However, we want to use this for vertex fetch, and data might not be aligned there, so handle it in this path too (basically just pass through alignment through to other functions). (It looks like it wouldn't work for for cached s3tc but this is no different than with AoS fetch.) Reviewed-by: Jose Fonseca <jfonseca@vmware.com>	2016-12-21 04:48:24 +01:00
Axel Davy	123e947228	st/nine: Upload on secondary context for DrawUp Avoid synchronization by using the secondary context for uploading the vertex data for DrawUp. v2: Rely on u_upload_mgr to use persistent coherent buffers. Do not flush. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:47:08 +01:00
Axel Davy	0ec4e5f630	st/nine: Dirty MANAGED buffers at Lock time Tests suggest MANAGED buffers are made dirty at Lock time, not at Unlock time. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:47:08 +01:00
Axel Davy	bad7f7cc63	st/nine: Implement new buffer upload path This new buffer upload path enables to lock faster than the normal path when using DISCARD/NOOVERWRITE. v2: Diverse cleanups and fixes. v3: Fix allocation size for 'lone' buffers and add more debug info. v4: Rewrite of the path to handle when DISCARD/NOOVERWRITE is not used anymore. The resource content is copied to the new resource used. v5: flush for safety after unmap (not sure it is really required here, but safer to flush). v6: Do not use the path if persistent coherent mapping is unavailable. Fix buffer creation flags. v7: Do not flush since it is not needed. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:47:08 +01:00
Axel Davy	8960be0e93	st/nine: Allow non-zero resource offset for vertex buffers Next patches will introduce an offset. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:47:08 +01:00
Axel Davy	1e64be6f91	st/nine: Do not wait for DEFAULT lock for volumes when we can If the volumes (and the texture container) are not referenced, then they are no pending operations on them. We can lock directly. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:47:08 +01:00
Axel Davy	b4f16615ef	st/nine: Do not wait for DEFAULT lock for surfaces when we can If the surfaces (and the texture container) are not referenced, then they are no pending operations on them. We can lock directly. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:47:08 +01:00
Axel Davy	525a1b292a	st/nine: Add arguments to context's blit and copy_region The new arguments enable to reference the objects while the function hasn't run. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:47:08 +01:00
Axel Davy	325324c749	st/nine: Idem for nine_context_gen_mipmap Will enable to use the bind count as an information for whether the surface/volume is used in the worker thread. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:47:08 +01:00
Axel Davy	7089d88199	st/nine: Bind destination for surface/volume uploads Will enable to use the bind count as an information for whether the surface/volume is used in the worker thread. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:47:08 +01:00
Axel Davy	d4a9b21feb	st/nine: Use nine_context_box_upload for volumes Use nine_context_box_upload for uploads: . systemmem volume to default volume . managed volume internal content to its resource. Check the uploads are executed before any action that can alter the data, that is LockBox and volume destruction. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:47:08 +01:00
Axel Davy	f042639231	st/nine: Fix leak with volume dtor The last level was not released. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:47:08 +01:00
Axel Davy	76e392d852	st/nine: Fix leak with cubetexture dtor The last level was not released. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:47:08 +01:00
Axel Davy	fec0b7f067	st/nine: Use nine_context_box_upload for surfaces Use nine_context_box_upload for uploads: . systemmem surface to default surface . managed surface internal content to its resource. Check the uploads are executed before any action that can alter the data, that is LockRect, NineSurface9_CopyDefaultToMem and surface destruction. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:47:08 +01:00
Axel Davy	c873a2bd0c	st/nine: Implement nine_context_box_upload This function will be used for surface and volume uploads Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:47:08 +01:00
Axel Davy	cadc7a5d94	st/nine: Use nine_context_gen_mipmap in BaseTexture9 Generate mipmaps in the worker thread. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:47:08 +01:00
Axel Davy	8d3e0f2187	st/nine: Implement nine_context_gen_mipmap To offload mipmap generation as well. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:47:08 +01:00
Axel Davy	16b6fb65ae	st/nine: Optimize managed buffer upload Do the upload in the other thread. Usually managed buffers are used once per frame. It is then very likely pending_upload is 0 at Lock time. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:47:08 +01:00
Axel Davy	a78b5f4378	st/nine: Implement nine_context_range_upload Will be used to upload buffers. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:47:08 +01:00
Axel Davy	1843e36b03	st/nine: Do not bind the container if forward is false This doesn't make sense to bind the container in that specific case. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:47:08 +01:00
Axel Davy	2fc8ef1401	st/nine: Comment and simplify iunknown The behaviour is a bit less obscure now. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:47:08 +01:00
Axel Davy	098ba64c4c	st/nine: Detach buffers in swapchain dtor. BackBuffers can survive swapchain dtor if the user has a reference on them. The swapchain itself has no reference on the buffer. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:47:08 +01:00
Axel Davy	14875ebd83	st/nine: Fix NineUnknown_Detach We don't bind the container in AddRef. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:47:08 +01:00
Axel Davy	930f479acf	st/nine: Simplify ARG_BIND_REF Remove some noop operations. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:47:08 +01:00
Axel Davy	9c4b4e8809	st/nine: Avoid flushing the queue for queries GetData Use the newly introduced counter to know when we don't need synchronization. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:47:08 +01:00
Patrick Rudolph	8a69343f1e	st/nine: Add CSMT_NO_WAIT_WITH_COUNTER Similar to the other macros, but introduces a counter, which enables to know when the instructions has been executed. Signed-off-by: Patrick Rudolph <siro@das-labor.org>	2016-12-20 23:47:08 +01:00
Axel Davy	884166a251	st/nine: Use nine_context_clear_render_target Enables to not wait for the worker thread for ColorFill. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:47:08 +01:00
Axel Davy	7b154ac04d	st/nine: Optimize ColorFill When we lock the whole surface to overwrite it, we can use DISCARD. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:23 +01:00
Axel Davy	9bf1da05d9	st/nine: Simplify ColorFill For render targets, NineSurface9_GetSurface is not expected to fail. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:23 +01:00
Axel Davy	31262bbce0	st/nine: use get_pipe_acquire/release when possible Use the acquire/release semantic when we don't need to wait for any pending command. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:23 +01:00
Axel Davy	22f6d6fbd2	st/nine: Implement Fast path for dynamic buffers and csmt Use the secondary pipe for DISCARD/NOOVERWRITE, which avoids stalling to get the pipe from the worker thread. v2: flush at unmap. This is required for example if the driver does hidden draw calls or copies. In the case of unsynchronized it is probably not required, but it is more safe. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:23 +01:00
Axel Davy	3e8234fff4	st/nine: Add secondary pipe for device The secondary pipe will be used for operations that don't need synchronization. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:23 +01:00
Axel Davy	7a7eeefd7d	st/nine: Add nine_context_get_pipe_acquire/release See commit for description. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:23 +01:00
Axel Davy	ddb6f1d2d1	st/nine: SYSTEMMEM ignores DISCARD. Tests show SYSTEMMEM should ignore DISCARD. Prevents game bugs with following patches reimplementing DISCARD. Halo is affected. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:23 +01:00
Axel Davy	4f344db8b0	st/nine: Upload Managed buffers just before draw call using them Previously we were uploading Managed buffers at the next draw call after they were set dirty. This is not the expected behaviour. Instead upload just before draw call needing the content. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:23 +01:00
Axel Davy	e52aded87f	st/nine: Track bindings for buffers Similar code than for textures. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:23 +01:00
Axel Davy	62068c9d90	st/nine: Fix BASETEX_REGISTER_UPDATE BASETEX_REGISTER_UPDATE was adding the texture to the list of textures to upload in too many cases. tex->base.base.bind will be set to true if the texture is in a stateblock, whereas we want to upload only if bound to the device, which is what bind_count is for. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:23 +01:00
Axel Davy	804b28cdc4	st/nine: Simplify the logic to bind textures This makes the code more readable. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:23 +01:00
Patrick Rudolph	fef23f6712	st/nine: Use nine_context for resource_copy_region Use nine_context wrapper for resource_copy_region. Enables to offload it with CSMT. Signed-off-by: Patrick Rudolph <siro@das-labor.org>	2016-12-20 23:44:23 +01:00
Patrick Rudolph	c8913a06b4	st/nine: Use nine_context for blit Enables to offload it with CSMT. Signed-off-by: Patrick Rudolph <siro@das-labor.org>	2016-12-20 23:44:23 +01:00
Patrick Rudolph	0fd5730613	st/nine: Add NINE_DEBUG=tid to turn threadid on or off To ease debugging. Signed-off-by: Patrick Rudolph <siro@das-labor.org>	2016-12-20 23:44:23 +01:00
Patrick Rudolph	3098bf03a6	st/nine: Print threadid in debug log To ease debugging. Signed-off-by: Patrick Rudolph <siro@das-labor.org>	2016-12-20 23:44:23 +01:00
Patrick Rudolph	ac2927335b	st/nine: Implement gallium nine CSMT Use an offloading thread for all nine_context functions. Macros are used to ease the reading of the code. Signed-off-by: Patrick Rudolph <siro@das-labor.org> Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:23 +01:00
Axel Davy	2c371a25a8	st/nine: Call GetPipe for implicit pipe usages With csmt, every usage of the pipe in the main thread has to be protected by calling GetPipe. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:23 +01:00
Patrick Rudolph	1277ceefd1	st/nine: Add struct nine_clipplane Required to know the size exact size of the plane. Signed-off-by: Patrick Rudolph <siro@das-labor.org>	2016-12-20 23:44:22 +01:00
Patrick Rudolph	3af17a671d	st/nine: Add nine_queue This queue mechanism will be used for CSMT. Signed-off-by: Patrick Rudolph <siro@das-labor.org>	2016-12-20 23:44:22 +01:00
Axel Davy	e068d3afe1	st/nine: Create pipe_surfaces on resource creation. Create the pipe_surfaces on renderable resources creation. This enables to avoid creating them on the fly. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	bb666b0297	st/nine: Back swvp in nine_context Part of the refactor to move all gallium calls to nine_state.c, and have all internal states required for those calls in nine_context. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	f5f881fd3e	st/nine: Change the way nine_shader gets the pipe The change is required with csmt, where depending on the thread you don't access the pipe the same way. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	97e4b65e7f	st/nine: Reimplement nine_context_apply_stateblock The new version uses nine_context functions instead of applying the changes directly to nine_context. This will enable it to work with CSMT. v2: Fix nine_context_light_enable_stateblock The memcpy arguments were wrong, and the state wasn't set dirty. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	8d967abb98	st/nine: Decompose nine_context_set_texture Part of the refactor to move all gallium calls to nine_state.c, and have all internal states required for those calls in nine_context. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	69f447752d	st/nine: Decompose nine_context_set_indices Part of the refactor to move all gallium calls to nine_state.c, and have all internal states required for those calls in nine_context. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	08b717dfd3	st/nine: Decompose nine_context_set_stream_source Part of the refactor to move all gallium calls to nine_state.c, and have all internal states required for those calls in nine_context. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	7ebdbb573b	st/nine: Do not use NineBaseTexture9 in nine_context Some fields are subject to modification outside of nine_context (SetLod, etc). Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	152d007769	st/nine: Move Managed Pool handling out of nine_context Part of the refactor to move all gallium calls to nine_state.c, and have all internal states required for those calls in nine_context. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	eb884a4ac2	st/nine: Integrate nine_pipe_context_clear to nine_context_clear Part of the refactor to move all gallium calls to nine_state.c, and have all internal states required for those calls in nine_context. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	b95205b1f2	st/nine: Move pipe and cso to nine_context Part of the refactor to move all gallium calls to nine_state.c, and have all internal states required for those calls in nine_context. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	66ad5b1592	st/nine: Rename pipe to pipe_data in nine_context This patch it to avoid name conflict when device->pipe will be moved to nine_context. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	fc49f7df89	st/nine: Rename cso in nine_context to cso_shader This patch it to avoid name conflict when device->cso is moved to nine_context. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	c7237e2c5c	st/nine: Access pipe_context via NineDevice9_GetPipe Except for nine_ff and nine_state. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	4a4eba8c05	st/nine: Remove NineDevice9_GetCSO Was useless. Remove useless usage in swapchain9. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	6a7541a5aa	st/nine: Move query9 pipe calls to nine_context This will enable to use threading for them. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	0a5252d25b	st/nine: Use atomics for nine_bind nine_bind didn't need atomics up to now, because it's use what always within a protected mutex. We need to use atomics because with the next patches several threads may use nine_bind. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	b748b8fd86	st/nine: Track dirty state groups in nine_context Part of the refactor to move all gallium calls to nine_state.c, and have all internal states required for those calls in nine_context. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	a0a18920c7	st/nine: Back User Clip Planes to nine_context Part of the refactor to move all gallium calls to nine_state.c, and have all internal states required for those calls in nine_context. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	c6ca7c747e	st/nine: Back ps to nine_context Part of the refactor to move all gallium calls to nine_state.c, and have all internal states required for those calls in nine_context. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	d671190df9	st/nine: Back ds to nine_context Part of the refactor to move all gallium calls to nine_state.c, and have all internal states required for those calls in nine_context. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	1a735a99d0	st/nine: Back all ff states in nine_context Part of the refactor to move all gallium calls to nine_state.c, and have all internal states required for those calls in nine_context. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	bb62ea925a	st/nine: Refactor LightEnable Call a helper function. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	cbe370020e	st/nine: Refactor SetLight Call a helper function to set the light. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	c5af96aebd	st/nine: Put ff data in a separate structure And make nine_state_access_transform take this new structure as input. Part of the refactor to move all gallium calls to nine_state.c, and have all internal states required for those calls in nine_context. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	4a6d83ebc2	st/nine: Back viewport to nine_context Part of the refactor to move all gallium calls to nine_state.c, and have all internal states required for those calls in nine_context. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	9498613607	st/nine: Back scissor to nine_context Part of the refactor to move all gallium calls to nine_state.c, and have all internal states required for those calls in nine_context. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	7f6e01052b	st/nine: Back RT to nine_context Part of the refactor to move all gallium calls to nine_state.c, and have all internal states required for those calls in nine_context. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	aafbd62955	st/nine: Back current index buffer to nine_context Part of the refactor to move all gallium calls to nine_state.c, and have all internal states required for those calls in nine_context. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	b13b217243	st/nine: Back all shader constants to nine_context For device vs shader float constants and may_swvp, the same tips than for the other constant types is used. Also memset the constants properly. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	93ac6dfdcc	st/nine: Back sampler states to nine_context Part of the refactor to move all gallium calls to nine_state.c, and have all internal states required for those calls in nine_context. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:22 +01:00
Axel Davy	2a698c3df2	st/nine: Back vs to nine_context And move programmable_vs storage and computation. Part of the refactor to move all gallium calls to nine_state.c, and have all internal states required for those calls in nine_context. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	43288cf376	st/nine: Back vdecl to nine_context Part of the refactor to move all gallium calls to nine_state.c, and have all internal states required for those calls in nine_context. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	63633e2a08	st/nine: Move stream freq data to nine_context Part of the refactor to move all gallium calls to nine_state.c, and have all internal states required for those calls in nine_context. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	848ffc81e4	st/nine: Move vtxbuf to nine_context Part of the refactor to move all gallium calls to nine_state.c, and have all internal states required for those calls in nine_context. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	aea7a019ef	st/nine: Move stream_usage_mask to nine_context Part of the refactor to move all gallium calls to nine_state.c, and have all internal states required for those calls in nine_context. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	eed47b748f	st/nine: Back textures into nine_context Part of the refactor to move all gallium calls to nine_state.c, and have all internal states required for those calls in nine_context. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	6bbb7b9fc5	st/nine: Move texture setting to nine_context_* And move samplers_shadow to nine_context. Part of the refactor to move all gallium calls to nine_state.c, and have all internal states required for those calls in nine_context. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	c1871e829a	st/nine: Track changed.texture only for stateblocks Part of the refactor to move all gallium calls to nine_state.c, and have all internal states required for those calls in nine_context. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	64e232bd60	st/nine: Move draw calls to nine_state Part of the refactor to move all gallium calls to nine_state.c, and have all internal states required for those calls in nine_context. v2: Release buffers for Draw*Up functions in device9.c, instead of nine_context. This prevents a leak with csmt where the wrong pointers were released. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	f72d8719eb	st/nine: Move core of device clear to nine_state Part of the refactor to move all gallium calls to nine_state.c, and have all internal states required for those calls in nine_context. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	1b24d5e1f5	st/nine: Introduce nine_context nine_context is a new structure which goal will be to contain all internal states. It will be the states of the second thread in the to-be-introduced CSMT mode. This patch moves several internal states to nine_context, while the next patches add the other fields. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	e3c59fbd25	st/nine: Implement WFOG properly We were advertising support for WFOG (like all win drivers), but we weren't implementing it. This patch implements the behaviour. See comments. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	b40f12ebf0	st/nine: Fix ff texture coordinate selection The code was wrongly detecting which texture coordinates to generate when the coordinate index was different to the stage index. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	c75415224f	st/nine: Convert redundant check to assert in ff ps We disable the alpha stage if the color stage is disabled. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	32f6f91617	st/nine: Fix two special cases in ff ps if first alpha stage is disabled and writes to temp, diffuse alpha is written to temp. Last stage always writes to current. Behaviour was deduced by tests with a test app. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	1efdc8f594	st/nine: Remove useless code in ff ps Current is already initialized to Diffuse. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	7afea63d4f	st/nine: Fix ff cases when stages should be disabled When a texture is read by a stage for colorop, it should be disabled, and disable following stages. When a texture is read for alphaop, 1.0f is read for the input, which is the behaviour for a dummy texture. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	191b90a35c	st/nine: Always initialize current in ff ps The check was not catching all possible cases. NVE4 should be fine. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	1ee978fa50	st/nine: Fix check for ff specular Fix the check for computing ff specular. This seems to match the opengl behavior, and give the correct output on windows. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	89716b0b38	st/nine: Do not saturate illumination coefficients in ff Fixes bad rendering of a test app. Wine has the same behaviour. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	877dc0308f	st/nine: Fix ff COLOR0 w component computation The computation was wrong. COLOR0's last component should be equal to the material diffuse w component. The behaviour was checked with a test app on Windows. Wine has the same behaviour. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	e94ac230a2	st/nine: Fix specular enable for alpha Apparently specular enable doesn't affect the alpha channel. Fixes https://github.com/iXit/Mesa-3D/issues/253 Behaviour comfirmed looking in wine sources. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	85811d0e87	st/nine: Ignore MULTISAMPLEMASK when RT is not multisampled We were ignoring MULTISAMPLEMASK for non-maskable multisample modes, but we were missing the non-multisampled case. Fixes a crash in Halo. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	bce9fe8db2	driconf: Fix missing gettext DRI_CONF_NINE_OVERRIDEVENDOR was missing gettext for the description. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	59048e7548	st/nine: Add new driconf options to control DISCARD behaviour See the patch for the new controls added. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	06657fa203	st/nine: Rework buffer presentation path Use the new API for DISCARD. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	35ea402a24	st/nine: Fix a leak in Swapchain dtor Count properly the number of backbuffers, and use the new info to release the correct number of buffers Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	b2f17e5f62	st/nine: Silent warnings with guid_str In non-debug build, the variables are unused, and thus trigger a compilation warning. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	2cd8622fb3	st/nine: Do not generate gallium NOP on d3d NOP Some drivers crash if NOP is generated. Besides there is no point to generate NOP. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	461e03167e	st/nine: Fix leak in user constant upload path The new code properly releases the previous buffers allocated. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	c3e5140142	st/nine: Correctly release sw cursor image cursor.image is used for software cursor emulation. It wasn't released. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	9c0f65e08a	st/nine: Handle when cursor stride is not what is expected SetCursor assumes for now a 32x32 argb cursor with pitch 128. 32x32 argb doesn't have pitch 128 on all hw, thus use a temporary surface with the correct pitch when needed. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:21 +01:00
Axel Davy	e7a0f580a6	st/nine: Avoid crash on empty Draw*Up Ignore empty draw calls. Avoid assertion fault when such draw calls happen in u_upload_mgr. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:20 +01:00
Axel Davy	ada0c2ceaa	st/nine: Capture texturestage states in pixel stateblocks pixels stateblocks need to capture these. Signed-off-by: Axel Davy <axel.davy@ens.fr>	2016-12-20 23:44:20 +01:00

... 2 3 4 5 6 ...

80504 Commits